Cancer risk prediction and evaluation system and evaluation method
Through the combination of specific SNP sites and primers combined with biochip technology, the existing cancer risk prediction process is solved, and the rapid and low-cost variety of cancer risk assessments are achieved, which improves detection efficiency and accuracy, and supports early screening and individualized treatment of cancer.
Patent Information
- Application Number
- CN202510498337.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-01
AI Technical Summary
The existing cancer risk prediction technology has problems such as cumbersome, time-consuming and high cost, which has affected the popularization of early screening and early diagnosis of cancer.
PCR amplification was performed using specific SNP sites and primer combinations, combined with biochip platform technology, cancer risk assessment was performed through fluorescence signal analysis, including primer combinations one and two, probe combinations and chip hybridization, and key SNP sites were screened using the GWAS database to establish a risk assessment system.
It has achieved rapid, low-cost, and multiple cancer risks assessment, improved detection efficiency and accuracy, reduced labor and reagent costs, and had high sensitivity and specificity to support early screening and individualized treatment of cancer.
Smart Images

Figure BDA0005367717000000031 
Figure BDA0005367717000000041 
Figure BDA0005367717000000051
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of cancer risk prediction, and particularly relates to a cancer risk prediction and assessment system and an assessment method. Background Art
[0002] Currently, the main techniques for cancer screening in clinical practice include liquid biopsy, gene detection, and biomarker detection. Some detection techniques still have problems such as insufficient sensitivity and specificity, and high costs.
[0003] Gene detection is a technique for detecting DNA molecules in the blood, body fluid cells of the tested person through methods such as gene chips, and analyzing the pathogenic genes, disease susceptibility genes, etc. contained in the tested person, and can detect whether there are abnormal genes, so as to predict the risk of cancer.
[0004] In current cancer risk prediction, gene detection technology still faces bottlenecks such as cumbersome processes, long time consumption, and high costs, which to a certain extent restricts the popularization process of early cancer screening and diagnosis. Summary of the Invention
[0005] In order to solve the problems in the prior art, the present invention provides a cancer risk prediction and assessment system and an assessment method, aiming to quickly, at low cost, and obtain the risk assessments of multiple cancers (lung cancer, esophageal cancer, gastric cancer, colorectal cancer, breast cancer) at one time.
[0006] The present invention solves its technical problems by adopting the following technical solutions:
[0007] The present invention aims to provide a cancer risk prediction and assessment system. The related SNP sites for cancer risk prediction include rs10937405, rs2131877, rs2736100, rs2853677, rs401681, rs7086803, rs748404, rs7741164, rs9387478, rs11066280, rs13042395, rs2274223, rs4072037, rs671, rs738722, rs10795668, rs3802842, rs6983267, rs7229639, rs961253, rs11615, rs1219648, rs2046210, rs2363956, rs3803662, rs3817198, rs4973768.
[0008] Further, it also includes primer combination one for amplifying SNP sites. The primer combination is as shown in SEQ ID NO.1-4, SEQ ID NO.7-10, SEQ ID NO.13-16, SEQ ID NO.19-22, SEQ ID NO.25-28, SEQ ID NO.31-34, SEQ ID NO.37-40, SEQ ID NO.43-46, SEQ ID NO.49-52, SEQ ID NO.55-58, SEQ ID NO.61-64, SEQ ID NO.67-70, SEQ ID NO.73-76, SEQ ID NO.79-82, SEQ ID NO.85-88, SEQ ID NO.91-94, SEQ ID NO.97-100, SEQ ID NO.103-106, SEQ ID NO.109-112, SEQ ID NO.115-118, SEQ ID NO.121-124, SEQ ID NO.127-130, SEQ ID NO.133-136, SEQ ID NO.139-142, SEQ ID NO.145-148, SEQ ID NO.151-154, SEQ ID NO.157-160 in the sequence listing.
[0009] Further, it also includes primer combination two for amplifying SNP sites. The primer combination is as shown in SEQ ID NO.163 and SEQ ID NO.164 in the sequence listing.
[0010] Further, it also includes a probe combination for chip hybridization. The probe combination is, for example, SEQ IQ NO.5, SEQ IQ NO.6, SEQ IQ NO.11, SEQ IQ NO.12, SEQ IQ NO.17, SEQ IQ NO.18, SEQ IQ NO.23, SEQ IQ NO.24, SEQ IQ NO.29, SEQ IQ NO.30, SEQ IQ NO.35, SEQ IQ NO.36, SEQ IQ NO.41, SEQ IQ NO.42, SEQ IQ NO.47, SEQ IQ NO.48, SEQ IQ NO.53, SEQ IQ NO.54, SEQ IQ NO.59, SEQ IQ NO.60, SEQ IQ NO.65, SEQ IQ NO.66, SEQ IQ NO.71, SEQ IQ NO.72, SEQ IQ NO.77, SEQ IQ NO.78, SEQ IQ NO.83, SEQ IQ NO.84, SEQ IQ NO.89, SEQ IQ NO.90, SEQ IQ NO.95, SEQ IQ NO.96, SEQ IQ NO.101, SEQ IQ NO.102, SEQ IQ NO.107, SEQ IQ NO.108, SEQ IQ NO.113, SEQ IQ NO.114, SEQ IQ NO.119, SEQ IQ NO.120, SEQ IQ NO.125, SEQ IQ NO.126, SEQ IQ NO.131, SEQ IQ NO.132, SEQ IQ NO.137, SEQ IQ NO.138, SEQ IQ NO.143, SEQ IQ NO.144, SEQ IQ NO.149, SEQ IQ NO.150, SEQ IQ NO.155, SEQ IQ NO.156, SEQ IQ NO.161, SEQ IQ NO.162 in the sequence listing.
[0011] An evaluation method for a cancer risk prediction and evaluation system includes the following steps:
[0012] The first-round amplification reaction: After mixing primer combination one, add Taq enzyme, dNTPs, PCR buffer, and gDNA sample, and perform the first-round amplification reaction in a PCR instrument.
[0013] The second-round amplification reaction: Mix the product of the first-round amplification, Taq enzyme, dNTPs, PCR buffer, and primer combination two, and perform the second-round amplification reaction in a PCR instrument.
[0014] Chip hybridization: After mixing the hybridization solution with the product of the second-round amplification reaction, it is dropped onto the chip spotting area fixed with the probe combination and incubated.
[0015] Data analysis: Scan with a scanner to read the fluorescence values at each locus on the gene chip. The fluorescence values are subjected to linear regression and slope correction with the "standard line" established using large sample data. After processing, the required values Cy3 and Cy5 are finally obtained, and the genotypes are classified.
[0016] Risk assessment: The OR values of the genotypes detected at all loci are multiplied and then divided by the number of loci to obtain the disease risk coefficient. If the disease risk coefficient is between 1 and 5, it is a low risk; between 5 and 10, it is a medium risk; and greater than 10, it is a high risk.
[0017] Furthermore, in the first-round amplification reaction, primer combination one is mixed to a final concentration of 50 μM, and the PCR instrument program is set as: 95°C, 2 minutes → 30 cycles (95°C, 30 seconds → 65°C, 3 minutes) → 4°C, ∞.
[0018] Furthermore, in the second-round amplification reaction, the PCR instrument program is set as: 94°C, 2 minutes → 35 cycles (95°C, 10 seconds → 65°C, 2 minutes) → 4°C, ∞.
[0019] Furthermore, in chip hybridization, the hybridization solution is SSC solution and Tris solution, and it is incubated for 6 hours under the reaction condition of 42°C.
[0020] Furthermore, the scanner uses a scanner equipped with 532 nm and 635 nm lasers. The fluorescence values include the values obtained by subtracting the background value from the foreground value of each chip locus in the 532 nm fluorescence band and the values obtained by subtracting the background value from the foreground value of each chip locus in the 635 nm fluorescence band.
[0021] Furthermore, the method for classifying genotypes includes: if Cy3 > 3000 and Cy5 < 3000, it is wild type; if Cy3 < 3000 and Cy5 > 3000, it is mutant type; if Cy3 > 3000 and Cy5 > 3000, it is heterozygous type.
[0022] The present invention designs a method for obtaining 27 SNP locus genes for risk prediction of multiple cancers (lung cancer, esophageal cancer, gastric cancer, colorectal cancer, breast cancer) and a risk assessment system. By accurately detecting specific SNP locus genes, combining biochip platform technology and a specific risk assessment system, it provides an efficient and accurate solution for early screening and risk assessment of cancers.
[0023] ① Screening and obtaining 27 key SNP loci by integrating GWAS database and scientific paper data. The specific information is as follows in Table 1:
[0024] Table 1
[0025]
[0026] Note: The three loci corresponding to gastric cancer, rs13042395, rs4072037, and rs738722, are the same as those in esophageal cancer, and the corresponding sequences in Table 2 below are also the same.
[0027] ②The present invention includes the following key systems:
[0028] Amplification System Kit: This system includes specific primer pairs designed for each cancer-related locus. The primer design takes into account the specificity of each SNP site and the length of the amplified product, enabling accurate amplification of gene fragments containing the target SNP site. Under the catalytic action of Taq enzyme, a large number of target DNA fragments are synthesized using dNTPs (deoxyribonucleotide triphosphates) as raw materials. Furthermore, to distinguish between wild-type and mutant forms, the 5' end of the universal primer sequence (primer combination 2) in the second round of amplification reaction is fluorescently labeled (universal primer 1 - Cy3, universal primer 2 - Cy5). This allows for better interpretation and analysis of test results through different fluorescence signal values (i.e., fluorescence numerical values).
[0029] The above-mentioned site-related primer sequences and probe sequences are shown in Table 2 below:
[0030] Table 2
[0031]
[0032]
[0033]
[0034]
[0035]
[0036]
[0037]
[0038]
[0039] Hybridization kit system: Probes that specifically bind to the target SNP sites were developed (as shown in Table 2). These probes are fluorescently labeled and immobilized on the chip surface to form a high-density microarray. The position of each probe is accurately marked on the chip to enable accurate identification and localization during detection. The hybridization process is carried out under specific temperature and time conditions to ensure the specific binding of the probes to the target DNA.
[0040] Data interpretation system: After the hybridization process is completed, a chip scanner is used to scan the biochip to obtain the fluorescence signal values (i.e., fluorescence numerical values) of each gene detection site. The fluorescence signal value is obtained by subtracting the background value from the foreground values of each chip site in two fluorescence bands, namely F532 - B and F635 - B. These numerical values are subjected to linear regression and slope correction with the "standard line" established using large sample data. After such processing, the required numerical values are finally obtained. The processed numerical values are classified as follows: If Cy3 > 3000 and Cy5 < 3000, it is wild type; if Cy3 < 3000 and Cy5 > 3000, it is mutant type; if Cy3 > 3000 and Cy5 > 3000, it is heterozygous type.
[0041] The log2 ratio value of each chip site is the logarithm to the base 2 of the value obtained by dividing F532 - B of each point by F635 - B. After dividing the log2 ratio numerical values of all SNP detection sites of the samples to be tested by the regression slope value, the slope - corrected ratio can be obtained. Then, after subtracting the linear regression intercept value from the slope - corrected ratio, the final ratio can be obtained. This final ratio is automatically aggregated, and the median of the final ratio numerical values of each of all samples to be tested after correction is taken. This median numerical value is the standard line of the log2 ratio numerical value of the SNP detection site.
[0042] Risk assessment system: First, a dedicated database is established by collecting, classifying, and summarizing the relevant disease site information in the GWAS database. The database contains the OR values of the case - study indicators for each site. The OR value of the wild - type is set to 1, the mutant type is OR * OR value, and the heterozygous type is OR value. Then, the disease risk coefficient is obtained by multiplying the OR values of the genotypes of all sites detected for the disease and then dividing by the number of sites. If the disease risk coefficient is between 1 and 5, it is low - risk; between 5 and 10, it is medium - risk; > 10, it is high - risk.
[0043] Compared with the prior art, the beneficial technical effects of the present invention are as follows:
[0044] 1. Compared with some traditional cancer gene detection methods, this invention offers certain advantages in terms of testing costs. It can simultaneously test multiple samples and multiple genes, thereby reducing the number of tests and the amount of reagents required, thus lowering testing costs. In addition, its relatively simple operation does not require complex manual analysis, which also reduces labor costs.
[0045] 2. The probes on the chip of the present invention have highly specific binding capabilities with target gene sequences, can accurately identify specific gene mutations to reduce misjudgments, and have high sensitivity and specificity, providing strong support for early diagnosis and prevention of cancer, and personalized treatment.
[0046] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above contents of the present invention and its objectives, features and advantages more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. DETAILED DESCRIPTION
[0047] The technical solutions of the present invention are further described in detail below with reference to specific embodiments. It should be understood that the following embodiments are merely exemplary illustrations and explanations of the present invention and should not be construed as limiting the scope of protection of the present invention. All technologies implemented based on the above content of the present invention are encompassed within the scope of protection that the present invention is intended to protect.
[0048] In addition, unless otherwise specified, various raw materials, reagents, instruments and equipment used in the present invention can be purchased from the market or prepared by existing methods.
[0049] Example 1
[0050] 1. Sample processing: Collect peripheral venous blood from the subject and extract gDNA samples according to the instructions for nucleic acid extraction or purification reagents.
[0051] 2. Implementation of the amplification system
[0052] ① First-round amplification: Mix the primers for the target loci in equal proportions to a final concentration of 50 μM. Add Taq enzyme, dNTPs, PCR buffer, and gDNA sample. Set the PCR instrument to the following program: 95°C, 2 minutes → 30 cycles (95°C, 30 seconds → 65°C, 3 minutes) → 4°C, ∞, and proceed with the amplification reaction.
[0053] ② Second round of amplification reaction: Mix the product of the first round of amplification, Taq enzyme, dNTPs, PCR buffer, and universal primers, and set the PCR instrument program to: 94°C, 2 minutes → 35 cycles (95°C, 10 seconds → 65°C, 2 minutes) → 4°C, ∞, and perform the second round of amplification reaction.
[0054] 3. Implementation of the hybridization system: After mixing the hybridization solution (SSC solution, Tris solution) with the product of the second round of amplification reaction, the mixture was evenly added dropwise to the sample spotting area of the chip, and then incubated at 42°C for 6 hours.
[0055] 4. Implementation of the cleaning system: First, add washing solution I (2×SSC, 0.2% SDS) to the cleaning box, then place the hybridized chip in it and wash it on a horizontal shaker (80 rpm) for 4 minutes; after washing, pour out the washing solution, add washing solution II (0.2×SSC) and wash under the same conditions for 4 minutes, finally centrifuge the chip to dry the surface liquid, and scan it with a scanner equipped with 532nm and 635nm lasers.
[0056] 5. Data Analysis: A scanner scans and reads the fluorescence values of each site on the gene chip at 532nm and 635nm laser wavelengths. The values are linearly regressed and slope corrected against a "standard line" established using large-sample data. After processing, the desired values are finally obtained. The processed values are classified as follows: if Cy3 > 3000 and Cy5 < 3000, it is wild-type; if Cy3 < 3000 and Cy5 > 3000, it is mutant; if Cy3 > 3000 and Cy5 > 3000, it is heterozygous.
[0057] 6. Results Analysis
[0058] ① Three Coriell standard DNA samples with known genotypes (NA20832, HG03166, and NA20801) were selected for testing. The genotype information is shown in Table 3.
[0059] Table 3
[0060] Locus ID Sample 1 (NA20832) Sample 2 (HG03166) Sample 3 (NA20801) rs10937405 CC TC CC rs2131877 AG AG GG rs2736100 AC CA CC rs2853677 AG GA AG rs401681 TT CT TT rs7086803 GG AG GG rs748404 TT TT CT rs7741164 GG GG GG rs9387478 AC AA CA rs11066280 TT TT TT rs13042395 CC CC CC rs2274223 AG GA AG rs4072037 CT TC CC rs671 GG GG GG rs738722 CT CC CT rs10795668 GG GG GG rs3802842 AC AA AC rs6983267 GT GG GT rs7229639 GG GA GG rs961253 AA CA CA rs11615 GA GG AG rs1219648 GG GA AA rs2046210 GA AA AA rs2363956 TG TT GG rs3803662 AG AA AG rs3817198 TT TT TT rs4973768 TT TC TT .
[0061] The experimental test results of the present invention are shown in Table 4 below.
[0062] Table 4
[0063]
[0064]
[0065] It has been verified that the sample detection results in this embodiment are consistent with the genotype results of the Coriell standard DNA sample, confirming that the detection kit of the present invention has a very high accuracy when applied to multiple SNP sites.
[0066] ②Set the OR value of the wild type to 1, the OR value of the mutant type to OR*OR value, and the OR value of the heterozygous type to OR value. The disease risk coefficient is obtained by multiplying the OR values of the genotypes detected at all loci of the disease and then dividing by the number of loci. If the disease risk coefficient is between 1 and 5, it is a low risk; between 5 and 10, it is a medium risk; and >10, it is a high risk, as shown in Table 5 below.
[0067] Table 5
[0068]
[0069]
[0070] The present invention detects genes and loci related to the susceptibility of cancers (lung cancer, esophageal cancer, gastric cancer, colorectal cancer, breast cancer), uses specific primers and probes in combination with microarray chip technology, and obtains genotyping data by analyzing the hybridization fluorescence signals at different positions on the chip. Then, a weighted genetic risk model is constructed based on a linear regression model, and finally, the model is used to perform individualized risk assessment on samples, which can simultaneously analyze the variations of several tumor susceptibility genes, significantly improving the detection efficiency and accuracy.
[0071] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0072] The above embodiments of the present invention have been described, but the present invention is not limited to the above specific embodiments. The above specific embodiments are only illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the purpose of the present invention and the scope protected by the claims, and these all belong to the protection scope of the present invention.
Claims
1. A cancer risk prediction and assessment system, characterized in that, The relevant SNP sites for cancer risk prediction include rs10937405, rs2131877, rs2736100, rs2853677, rs401681, rs7086803, rs748404, rs7741164, rs9387478, rs11066280, rs13042395, rs2274223, rs4072037, rs671, rs738722, rs10795668, rs3802842, rs6983267, rs7229639, rs961253, rs11615, rs1219648, rs2046210, rs2363956, rs3803662, rs3817198, rs4973768.
2. The cancer risk prediction and assessment system according to claim 1, characterized in that: It also includes primer combination one for amplifying SNP sites. Primer combination one is as shown in SEQ IQ NO.1-4, SEQ IQ NO.7-10, SEQ IQ NO.13-16, SEQ IQ NO.19-22, SEQ IQ NO.25-28, SEQ IQ NO.31-34, SEQ IQ NO.37-40, SEQ IQ NO.43-46, SEQ IQ NO.49-52, SEQ IQ NO.55-58, SEQ IQ NO.61-64, SEQ IQ NO.67-70, SEQ IQNO.73-76, SEQ IQ NO.79-82, SEQ IQ NO.85-88, SEQ IQ NO.91-94, SEQ IQ NO.97-100, SEQIQ NO.103-106, SEQ IQ NO.109-112, SEQ IQ NO.115-118, SEQ IQ NO.121-124, SEQ IQNO.127-130, SEQ IQ NO.133-136, SEQ IQ NO.139-142, SEQ IQ NO.145-148, SEQ IQNO.151-154, SEQ IQ NO.157-160 in the sequence listing.
3. The cancer risk prediction and assessment system according to claim 2, characterized in that: It also includes primer combination two for amplifying SNP sites. Primer combination two is as shown in SEQ IQ NO.163 and SEQ IQ NO.164 in the sequence listing.
4. The cancer risk prediction and assessment system according to claim 3, wherein: It also includes a probe combination for chip hybridization, and the probe combination is as shown in SEQ ID NO.5, SEQ ID NO.6, SEQ ID NO.11, SEQ ID NO.12, SEQ ID NO.17, SEQ ID NO.18, SEQ ID NO.23, SEQ ID NO.24, SEQ ID NO.29, SEQ ID NO.30, SEQ ID NO.35, SEQ ID NO.36, SEQ ID NO.41, SEQ ID NO.42, SEQ ID NO.47, SEQ ID NO.48, SEQ ID NO.53, SEQ ID NO.54, SEQ ID NO.59, SEQ ID NO.60, SEQ ID NO.65, SEQ ID NO.66, SEQ ID NO.71, SEQ ID NO.72, SEQ ID NO.77, SEQ ID NO.78, SEQ ID NO.83, SEQ ID NO.84, SEQ ID NO.89, SEQ ID NO.90, SEQ ID NO.95, SEQ ID NO.96, SEQ ID NO.101, SEQ ID NO.102, SEQ ID NO.107, SEQ ID NO.108, SEQ ID NO.113, SEQ ID NO.114, SEQ ID NO.119, SEQ ID NO.120, SEQ ID NO.125, SEQ ID NO.126, SEQ ID NO.131, SEQ ID NO.132, SEQ ID NO.137, SEQ ID NO.138, SEQ ID NO.143, SEQ ID NO.144, SEQ ID NO.149, SEQ ID NO.150, SEQ ID NO.155, SEQ ID NO.156, SEQ ID NO.161, SEQ ID NO.162 in the sequence listing.
5. The evaluation method of a cancer risk prediction and assessment system according to any one of claims 1-4, characterized in that: It includes the following steps: The first round of amplification reaction: After the primer combination one is mixed, Taq enzyme, dNTPs, PCR buffer, and gDNA sample are added, and the first round of amplification reaction is carried out in a PCR instrument; The second round of amplification reaction: The product of the first round of amplification, Taq enzyme, dNTPs, PCR buffer, and primer combination two are mixed, and the second round of amplification reaction is carried out in a PCR instrument; Chip hybridization: After the hybridization solution is mixed with the product of the second round of amplification reaction, it is dropped onto the chip spotting area fixed with the probe combination for incubation; Data analysis: Scan with a scanner to read the fluorescence values at each site on the gene chip. The fluorescence values are subjected to linear regression and slope correction with the "standard line" established using large sample data. After processing, the required values Cy3 and Cy5 are finally obtained, and the genotypes are classified. Risk assessment: The OR values of the genotypes detected at all sites are multiplied and then divided by the number of sites to obtain the disease risk coefficient. If the disease risk coefficient is between 1 and 5, it is a low risk; between 5 and 10, it is a medium risk; and greater than 10, it is a high risk.
6. The evaluation method of a cancer risk prediction and evaluation system according to claim 5, characterized in that: In the first-round amplification reaction, primer combination one was mixed to a final concentration of 50 μM, and the PCR instrument program was set as: 95°C, 2 minutes → 30 cycles (95°C, 30 seconds → 65°C, 3 minutes) → 4°C, ∞.
7. The evaluation method of a cancer risk prediction and evaluation system according to claim 5, characterized in that: In the second-round amplification reaction, the PCR instrument program was set as: 94°C, 2 minutes → 35 cycles (95°C, 10 seconds → 65°C, 2 minutes) → 4°C, ∞.
8. The evaluation method of a cancer risk prediction and evaluation system according to claim 5, characterized in that: In chip hybridization, the hybridization solution is SSC solution and Tris solution, and it is incubated for 6 hours under the reaction condition of 42°C.
9. The evaluation method of a cancer risk prediction and evaluation system according to claim 5, characterized in that: The scanner uses a scanner equipped with 532 nm and 635 nm lasers. The fluorescence values include the values obtained by subtracting the background value from the foreground value of each chip site in the 532 nm fluorescence band and the values obtained by subtracting the background value from the foreground value of each chip site in the 635 nm fluorescence band.
10. The evaluation method of a cancer risk prediction and assessment system according to claim 9, characterized in that: The method for classifying genotypes includes: If Cy3 > 3,000 and Cy5 < 3,000, it is wild type; if Cy3 < 3,000 and Cy5 > 3,000, it is mutant type; if Cy3 > 3,000 and Cy5 > 3,000, it is heterozygous type.