Pancreatic cancer risk estimation method and system based on pancreatic cancer related genome markers
By constructing a risk prediction model based on multiple single nucleotide polymorphism sites and rare germline mutation genes associated with pancreatic cancer, and combining artificial intelligence algorithms and multi-model evaluation strategies, the shortcomings of existing pancreatic cancer risk assessment technologies have been addressed, achieving stable and accurate risk prediction and resource optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AFFILIATED HUSN HOSPITAL OF FUDAN UNIV
- Filing Date
- 2025-12-26
- Publication Date
- 2026-05-05
AI Technical Summary
There is a lack of effective pancreatic cancer risk prediction models in the current technology, and conventional genomic biomarker studies have small sample sizes, low risk values, and low germline mutation rates, making it difficult to achieve accurate risk assessment.
We constructed a risk prediction model based on multiple single nucleotide polymorphism sites and rare germline mutations associated with pancreatic cancer. Combining artificial intelligence algorithms, we conducted pancreatic cancer risk assessment through data hierarchical processing and multi-model evaluation strategies. This included collecting basic information, gene testing, data encoding and preprocessing, risk feature set formation, risk assessment, and intervention recommendations.
It enables quantitative prediction of the lifetime risk of pancreatic cancer, improves the stability and applicability of risk assessment, reduces the risk of over-testing and under-testing, optimizes the allocation of medical resources, and has good prospects for clinical application.
Smart Images

Figure CN121983120A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of pancreatic cancer genomics analysis, specifically to a method and system for pancreatic cancer risk prediction based on pancreatic cancer-related genomic biomarkers. Background Technology
[0002] Pancreatic cancer is a common digestive tract tumor. Pancreatic ductal adenocarcinoma is the main pathological type of pancreatic cancer, accounting for 85%-95% of all pancreatic cancers (hereinafter, all pancreatic cancers refer to pancreatic ductal adenocarcinoma). According to data from the World Health Organization's 2018 Global Cancer Watch, there were 458,918 new cases of pancreatic cancer globally, resulting in 432,242 deaths, making it the seventh leading cause of cancer death for both men and women. According to the latest 2020 projections from the American Cancer Society, the number of new pancreatic cancer cases in the United States in 2020 was expected to reach 57,600, with 47,050 deaths. The number of deaths is close to the number of new cases. Compared to developed countries, the incidence and mortality rates of pancreatic cancer are higher in developing countries. According to 2015 cancer data in my country, there were 90,100 new cases of pancreatic cancer that year, with 79,400 deaths. The mortality rate remains high, with an average 5-year survival rate of only about 6%. Pancreatic cancer is a major disease that poses a significant threat to the health and lives of the people, as it imposes a heavy burden on patients' quality of life.
[0003] With the implementation of the Human Genome Project (HGP), the establishment of the HapMap database, and the rapid development of next-generation sequencing technology, genomic research has begun to demonstrate its important role in cancer research. Various genetic markers not only serve as diagnostic markers but also provide new research methods and opportunities for studying the molecular mechanisms of pancreatic cancer, searching for novel pancreatic cancer-specific tumor markers, and elucidating the pathogenesis of pancreatic cancer. Genomic research includes single nucleotide polymorphisms (SNPs), copy number variations (CNVs), and germline mutations.
[0004] Genomic biomarkers, as stably inherited markers of the human genome, can be directly detected through peripheral blood samples. Genomic biomarkers proven to be associated with diseases can be used for disease risk assessment. Therefore, using genomic biomarkers for pancreatic cancer risk prediction is entirely feasible.
[0005] However, conventional studies on pancreatic cancer genomic biomarkers have faced practical challenges, including limited sample sizes, lower risk values for genomic biomarkers compared to other tumors, and lower germline mutation rates. With the support of artificial intelligence algorithms, it is possible to construct a pancreatic cancer risk prediction model that can be practically applied in clinical practice. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for pancreatic cancer risk prediction based on pancreatic cancer-related genomic biomarkers. The aim is to provide an artificial intelligence-assisted risk prediction model for pancreatic cancer-related genomic biomarkers, thus solving the problem of the lack of risk prediction models for pancreatic cancer.
[0007] To achieve the above objectives, the present invention provides a method for pancreatic cancer risk prediction based on pancreatic cancer-related genomic biomarkers, comprising the following steps:
[0008] S1. Collect basic information about the subject to be evaluated. The basic information includes at least gender, age, and whether the subject has a history of cancer or a family history of cancer.
[0009] S2, collect peripheral blood samples from the subject to be evaluated, and perform gene testing on the peripheral blood samples to obtain genomic biomarker data related to pancreatic cancer;
[0010] S3, the genomic biomarker data and the basic information are encoded and preprocessed, and the data is filtered and weighted based on a pancreatic cancer-specific data screening algorithm to form a pancreatic cancer risk feature set;
[0011] S4, input the pancreatic cancer risk feature set into the pre-constructed pancreatic cancer risk prediction model to obtain the pancreatic cancer risk assessment result of the subject to be assessed;
[0012] S5. Calculate the pancreatic cancer risk ratio based on the risk assessment results, and stratify the subjects to be assessed according to a preset threshold.
[0013] S6. Output corresponding health management or medical intervention recommendations based on the risk stratification results.
[0014] Furthermore, the pancreatic cancer-related genomic biomarkers in step S2 include single nucleotide polymorphism (SNP) detection sites, which include at least one or more of the following sites:
[0015] rs401681, rs2255280, rs2317900, rs6971499, rs167020, rs1561927, rs9581943, r s4885093, rs9543325, rs7190458, rs372883, rs1547374, rs16986825, rs5768709.
[0016] Furthermore, the pancreatic cancer-related genomic markers also include rare germline mutant genes, which include at least one or more of the following genes:
[0017] AR, ATM, ATR, BARD1, BRAF, BRCA1, BRCA2, BRIP1, CDH1, CDK12, CHEK1, CHEK2, ERBB2, ESR1, FANCA, FANCL, HDAC2, HOXB13, KRAS, MRE11, NBN, NRAS, PALB2, PIK3CA, PPP2R2A, PTEN, RAD51B, RAD51C, RAD51D, RAD54L, STK11, TP53.
[0018] Furthermore, the pancreatic cancer-specific data screening algorithm in step S3 includes:
[0019] The genomic biomarker data and basic information were mapped according to their roles in the pathogenesis of pancreatic cancer into a common genetic susceptibility layer, a rare mutation layer, and a clinical regulatory layer, wherein:
[0020] Common genetic susceptibility layers are used to characterize the cumulative genetic susceptibility of multiple single nucleotide polymorphism sites;
[0021] Rare mutation layers are used to characterize the risk of high-impact pancreatic cancer-related germline mutations;
[0022] The clinical regulatory layer is used to regulate genetic risk in relation to age and cancer history.
[0023] Furthermore, in the aforementioned common genetic susceptibility layer, the genotypes of single nucleotide polymorphism sites are encoded and weighted as follows:
[0024] First, each locus is coded hierarchically based on its allele status to reflect the cumulative number of mutated alleles.
[0025] Then, based on the stability characteristics of the loci in the detection population, a corresponding confidence correction factor is assigned to each locus;
[0026] The hierarchical coding results are combined with the corresponding confidence correction factor to reduce the impact of detecting highly volatile sites on the overall risk assessment.
[0027] Furthermore, in the rare mutation layer, rare germline mutations are processed according to the following rules:
[0028] When no rare germline mutations associated with pancreatic cancer are detected, the rare mutation layer is marked as a low-risk state;
[0029] When at least one rare germline mutation associated with pancreatic cancer is detected, the rare mutation layer is uniformly mapped to a single high-risk marker;
[0030] This is to avoid the unexpected cumulative amplification effect of multiple rare germline mutations during the risk assessment process.
[0031] Furthermore, in step S3, the data encoding and preprocessing at least includes:
[0032] For missing genotype data, representative genotypes based on population distribution characteristics are used as replacements.
[0033] Age information is centralized and scaled uniformly to eliminate the impact of different dimensions on risk assessment;
[0034] Information on past cancer history or family cancer history should be recorded using a yes / no statement.
[0035] Furthermore, in step S4, the pancreatic cancer risk prediction model is constructed based on a multi-model evaluation strategy, and the model includes at least one of a logistic regression model, a support vector machine model, a random forest model, or a gradient boosting tree model.
[0036] On the other hand, the present invention also provides a pancreatic cancer risk prediction system based on pancreatic cancer-related genomic biomarkers, comprising:
[0037] The basic information collection module is used to obtain information such as the gender, age, and past or family history of cancer of the subject to be evaluated.
[0038] The gene detection data acquisition module is used to acquire genomic marker detection data from the peripheral blood sample of the subject to be evaluated.
[0039] The data encoding and preprocessing module is used to encode genomic biomarker data and basic information, handle missing values, and standardize them.
[0040] A pancreatic cancer-specific data screening module is used to stratify, screen, and weight the data according to the relevant genetic mechanisms of pancreatic cancer, and generate a pancreatic cancer risk feature set.
[0041] The risk prediction model module is used to output pancreatic cancer risk assessment results based on the pancreatic cancer risk feature set;
[0042] The risk assessment and stratification module is used to calculate the risk ratio for pancreatic cancer and classify risk levels.
[0043] The intervention recommendation output module is used to output corresponding health management or medical intervention recommendations based on the risk level;
[0044] The modules work together to implement the method described in claim 1.
[0045] On the other hand, the present invention also provides a computer-readable storage medium, characterized in that the storage medium stores a plurality of instructions adapted for loading by a processor to execute the steps in the above method.
[0046] This invention has the following technological advancements and beneficial effects:
[0047] 1. For the first time, a systematic risk prediction based on multiple pancreatic cancer-related genomic biomarkers was achieved.
[0048] This invention does not rely on a single high-risk mutation or a single site to determine the risk of pancreatic cancer. Instead, it comprehensively utilizes multiple single nucleotide polymorphism sites and rare germline mutation genes associated with pancreatic cancer to construct a risk feature set that reflects the genetic susceptibility of the population and high-risk events in individuals. This enables a quantitative prediction of the lifetime risk of pancreatic cancer, filling the gap in the existing technology for pancreatic cancer risk prediction models.
[0049] 2. Introduce pancreatic cancer-specific data screening and risk expression rules to improve the stability of risk assessment.
[0050] This invention addresses the characteristics of pancreatic cancer, such as dispersed genetic features and limited contribution from a single locus. It performs stratified processing of gene detection data and avoids abnormal amplification effects of individual loci or a few samples on risk assessment results through stability correction and rare mutation risk saturation. This results in better stability and reproducibility of risk assessment results in different populations and under different testing conditions.
[0051] 3. By combining clinical regulatory information, risk stratification can be achieved that more closely reflects the actual population distribution.
[0052] By incorporating age and past or family history of cancer as moderating factors into the risk assessment process, this invention avoids the limitations of mechanical judgment based solely on genetic data, making the risk stratification results more consistent with the statistical characteristics of pancreatic cancer occurrence in the actual population.
[0053] 4. Employ a multi-model evaluation strategy to adapt to the needs of different application scenarios.
[0054] This invention does not limit itself to a single risk assessment model, but rather uses a multi-model assessment strategy to flexibly select model types in different application scenarios such as population screening and high-risk population assessment, thereby achieving a balance between specificity and sensitivity and reducing the risk of over-testing and missed detection.
[0055] 5. It has clear medical application value and significance for promotion.
[0056] By avoiding unnecessary pancreatic examinations for low-risk individuals and providing advance risk warnings and monitoring for high-risk individuals, this invention helps optimize the allocation of medical resources, reduce the medical burden on individuals and society, and has good clinical application prospects and promotional value. Attached Figure Description
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0058] Figure 1 This is a flowchart of a pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers, according to Embodiment 1 of the present invention.
[0059] Figure 2 This is a system framework diagram of a pancreatic cancer risk prediction system based on pancreatic cancer-related genomic biomarkers, according to Embodiment 2 of the present invention. Detailed Implementation
[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0061] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0062] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0063] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0064] This invention relates to a method and system for pancreatic cancer risk prediction based on pancreatic cancer-related genomic biomarkers. To better illustrate the technical solution of this invention, specific embodiments are shown below.
[0065] Example 1
[0066] This embodiment provides a method for pancreatic cancer risk prediction based on pancreatic cancer-related genomic biomarkers, specifically including the following steps:
[0067] S1, Collect basic information about the object to be evaluated.
[0068] In practical applications, basic information includes at least gender, age, and whether there is a past or family history of cancer. This information is obtained through manual entry, electronic medical record interfaces, or automatic acquisition by health management systems, and is stored in structured field format. This basic information is not directly used as a risk conclusion in subsequent processing, but rather to adjust for genetic risk outcomes to avoid risk judgment bias caused by relying solely on genetic data.
[0069] For example, the basic information collected for a subject to be evaluated is: male, 52 years old, with no clear history of cancer or family history of cancer.
[0070] S2, collect peripheral blood samples and obtain data on pancreatic cancer-related genomic biomarkers.
[0071] In this step, peripheral blood samples are collected from the subject of evaluation. The sample volume can be 3–10 mL, preferably 5 mL. After nucleic acid extraction, the peripheral blood samples are used to obtain genomic biomarker data related to pancreatic cancer through genotyping or sequencing. The genomic biomarker data includes at least two types:
[0072] One type consists of genotype data with multiple single nucleotide polymorphism sites.
[0073] Another category is the detection results of rare germline mutation genes associated with pancreatic cancer.
[0074] For example, in one test, the following partial results were obtained:
[0075] rs401681 is heterozygous, rs2255280 is homozygous wild-type, rs2317900 is heterozygous, rs6971499 is homozygous mutant, and rs1561927 failed detection.
[0076] Meanwhile, no rare germline mutations associated with pancreatic cancer, such as BRCA1, BRCA2, PALB2, and STK11, were detected.
[0077] S3 encodes and preprocesses genomic biomarker data and basic information to form a pancreatic cancer risk feature set.
[0078] In this step, data from different sources are first organized and standardized to serve as effective input for subsequent risk assessment. To reflect the genetic risk characteristics of pancreatic cancer, the data are divided into three categories: information related to common genetic susceptibility, information related to rare mutations, and information related to clinical regulation.
[0079] For single nucleotide polymorphism (SNP) sites, each site is graded according to the detected allele status to reflect the accumulation of mutant alleles. When the detection result for a certain site is missing, the site is not discarded directly. Instead, a representative genotype is selected as a substitute value based on the distribution characteristics of the site in the population, thereby avoiding the imbalance of overall characteristics caused by missing data.
[0080] After completing the hierarchical expression, each site is not directly used with equal weight. Instead, different levels of correction are applied to different sites based on their stability characteristics during the detection process. Sites with higher stability and better detection repeatability are given higher weight in the risk characteristics, while the impact of sites with greater detection volatility on the final risk assessment is correspondingly weakened.
[0081] For rare germline mutation information, this embodiment does not distinguish the specific number or type of mutation, but focuses on the presence of rare mutations highly associated with pancreatic cancer. When at least one relevant rare germline mutation is detected, the dimension is considered to have a significant genetic risk, and this information is mapped to a uniform high-risk marker; when no relevant mutation is detected, the dimension is marked as low-risk. In this way, the overall assessment results are not abnormally amplified due to a large number of rare mutations in individual samples.
[0082] Age data in the basic information is scaled uniformly before entering the risk assessment so that data of different dimensions can participate in the calculation together with genetic information; information on past cancer history or family cancer history is recorded in the form of whether it exists, and is used to modulate the genetic risk outcome.
[0083] After completing the above processing, common genetic susceptibility information, rare mutation information, and clinical regulatory information are combined to form a risk feature set for pancreatic cancer risk assessment.
[0084] In a specific example, for the aforementioned subject A to be evaluated, its risk feature set may include the following: multiple single nucleotide polymorphism site feature values after graded expression and stability correction, such as the corrected feature value corresponding to rs401681, the corrected feature value corresponding to rs2317900, and the corrected feature value corresponding to rs6971499; the overall risk marker value corresponding to the rare mutation layer, used to indicate whether there are rare germline mutations related to pancreatic cancer; and the age feature value and the regulatory feature value corresponding to the past history of tumors or family history of tumors after scale normalization.
[0085] The aforementioned risk feature set can be represented in vector or tabular form, where each feature corresponds to a processed genetic or clinical piece of information, characterizing the individual's overall status in terms of pancreatic cancer-related genetic susceptibility, rare high-risk events, and clinical moderating factors. This risk feature set is used as the overall input for subsequent pancreatic cancer risk assessment calculations, rather than making independent risk judgments on individual features.
[0086] S4 obtains pancreatic cancer risk assessment results based on a multi-model assessment strategy.
[0087] In this step, the risk feature set formed in step S3 is input into the pancreatic cancer risk prediction model for calculation. The risk prediction model is not a single model, but is constructed based on a multi-model evaluation strategy, including at least one of the following: logistic regression model, support vector machine model, and tree-structured ensemble model.
[0088] During model building, multiple rounds of training and validation on historical sample data ensure that the model maintains stable risk discrimination capabilities under different sample distribution conditions. Depending on the specific application requirements, different models can be selected for risk assessment: in scenarios emphasizing overall screening stability, models insensitive to outliers are preferred; in scenarios emphasizing high-risk identification capabilities, models more sensitive to nonlinear features are preferred.
[0089] For example, in a population screening application, after inputting the risk feature set of a subject to be evaluated into a logistic regression model, the corresponding pancreatic cancer risk ratio was obtained as 0.82.
[0090] S5, risk stratification based on risk assessment results.
[0091] After obtaining the risk ratio, it is compared with a preset threshold to stratify the subjects to be assessed based on risk. Subjects with a risk ratio below the first threshold are classified as low-risk, those with a risk ratio between the first and second thresholds are classified as medium-risk, and those with a risk ratio above the second threshold are classified as high-risk.
[0092] In the example above, after inputting the risk feature set of the subjects to be evaluated into the pancreatic cancer risk prediction model, a pancreatic cancer risk ratio of 0.82 was obtained. According to the preset risk stratification rules, subjects with a risk ratio less than 1 were classified as low-risk. Based on this, the system outputs the corresponding risk assessment results and health management recommendations. S6, Output risk assessment results and intervention recommendations.
[0093] Based on the risk stratification results, the system generates corresponding prompts and health management recommendations. For low-risk individuals, the system indicates that no additional pancreatic examination is required; for medium- and high-risk individuals, the system suggests increasing the frequency of physical examinations and, if necessary, conducting pancreatic-related imaging examinations or further testing.
[0094] S6. Output corresponding health management or medical intervention recommendations based on the risk stratification results.
[0095] Example 2
[0096] This embodiment provides a pancreatic cancer risk prediction system for implementing the above-described method. The system includes a basic information acquisition module 1, a gene detection data acquisition module 2, a data encoding and preprocessing module 3, a pancreatic cancer-specific data screening module 4, a risk prediction model module 5, a risk assessment and stratification module 6, and an intervention suggestion output module 7.
[0097] The system comprises the following modules: Basic Information Acquisition Module (BIA) for acquiring and organizing information on gender, age, and tumor history; Gene Testing Data Acquisition Module (GTA) for receiving gene testing results from peripheral blood samples; Data Encoding and Preprocessing Module (GC&P) for standardizing formats, handling missing data, and scaling data from different sources; Pancreatic Cancer-Specific Data Screening Module (PCS) for performing genetic information stratification, site correction, and rare mutation handling; Risk Prediction Model Module (GPMM) for outputting risk assessment results based on risk feature sets; Risk Assessment and Stratification Module (SAT) for determining risk levels; and Intervention Recommendation Output Module (IRO) for generating health management recommendations corresponding to the risk level.
[0098] Example 3
[0099] This embodiment provides a computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program is used to implement steps S1 to S6 in the above-described pancreatic cancer risk prediction method.
[0100] The above embodiments describe specific implementations of the present invention, including the experimental steps and data analysis process of the peripheral blood immune cell atlas construction method, and demonstrate the implementation details of the system and device. These implementations can be optimized or adjusted according to specific needs to adapt to different experimental scenarios or clinical applications.
[0101] Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting pancreatic cancer risk based on pancreatic cancer-related genomic biomarkers, characterized in that, Includes the following steps: S1. Collect basic information about the subject to be evaluated. The basic information includes at least gender, age, and whether the subject has a history of cancer or a family history of cancer. S2, collect peripheral blood samples from the subject to be evaluated, and perform gene testing on the peripheral blood samples to obtain genomic biomarker data related to pancreatic cancer; S3, the genomic biomarker data and the basic information are encoded and preprocessed, and the data is filtered and weighted based on a pancreatic cancer-specific data screening algorithm to form a pancreatic cancer risk feature set; S4, input the pancreatic cancer risk feature set into the pre-constructed pancreatic cancer risk prediction model to obtain the pancreatic cancer risk assessment result of the subject to be assessed; S5. Calculate the pancreatic cancer risk ratio based on the risk assessment results, and stratify the subjects to be assessed according to a preset threshold. S6. Output corresponding health management or medical intervention recommendations based on the risk stratification results.
2. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 1, characterized in that, The pancreatic cancer-related genomic biomarkers in step S2 include single nucleotide polymorphism (SNP) detection sites, which include at least one or more of the following sites: rs401681, rs2255280, rs2317900, rs6971499, rs167020, rs1561927, rs9581943, r s4885093, rs9543325, rs7190458, rs372883, rs1547374, rs16986825, rs5768709.
3. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 1 or 2, characterized in that, The pancreatic cancer-related genomic markers also include rare germline mutant genes, which include at least one or more of the following genes: AR, ATM, ATR, BARD1, BRAF, BRCA1, BRCA2, BRIP1, CDH1, CDK12, CHEK1, CHEK2, ERBB2, ESR1, FANCA, FANCL, HDAC2, HOXB13, KRAS, MRE11, NBN, NRAS, PALB2, PIK3CA, PPP2R2A, PTEN, RAD51B, RAD51C, RAD51D, RAD54L, STK11, TP53.
4. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 1, characterized in that, The pancreatic cancer-specific data screening algorithm described in step S3 includes: The genomic biomarker data and basic information were mapped according to their roles in the pathogenesis of pancreatic cancer into a common genetic susceptibility layer, a rare mutation layer, and a clinical regulatory layer, wherein: Common genetic susceptibility layers are used to characterize the cumulative genetic susceptibility of multiple single nucleotide polymorphism sites; Rare mutation layers are used to characterize the risk of high-impact pancreatic cancer-related germline mutations; The clinical regulatory layer is used to regulate genetic risk in relation to age and cancer history.
5. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 4, characterized in that, In the aforementioned common genetic susceptibility layers, the genotypes of single nucleotide polymorphism sites are encoded and weighted as follows: First, each locus is coded hierarchically based on its allele status to reflect the cumulative number of mutated alleles. Then, based on the stability characteristics of the loci in the detection population, a corresponding confidence correction factor is assigned to each locus; The hierarchical coding results are combined with the corresponding confidence correction factor to reduce the impact of detecting highly volatile sites on the overall risk assessment.
6. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 4, characterized in that, In the rare mutation layer, rare germline mutations are handled according to the following rules: When no rare germline mutations associated with pancreatic cancer are detected, the rare mutation layer is marked as a low-risk state; When at least one rare germline mutation associated with pancreatic cancer is detected, the rare mutation layer is uniformly mapped to a single high-risk marker; This is to avoid the unexpected cumulative amplification effect of multiple rare germline mutations during the risk assessment process.
7. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 1, characterized in that, In step S3, the data encoding and preprocessing includes at least: For missing genotype data, representative genotypes based on population distribution characteristics are used as replacements. Age information is centralized and scaled uniformly to eliminate the impact of different dimensions on risk assessment; Information on past cancer history or family cancer history should be recorded using a yes / no statement.
8. The pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in claim 1, characterized in that, In step S4, the pancreatic cancer risk prediction model is constructed based on a multi-model evaluation strategy, and the model includes at least one of the following: logistic regression model, support vector machine model, random forest model, or gradient boosting tree model.
9. A pancreatic cancer risk prediction system based on pancreatic cancer-related genomic biomarkers, characterized in that, include: The basic information collection module is used to obtain information such as the gender, age, and past or family history of cancer of the subject to be evaluated. The gene detection data acquisition module is used to acquire genomic marker detection data from the peripheral blood sample of the subject to be evaluated. The data encoding and preprocessing module is used to encode genomic biomarker data and basic information, handle missing values, and standardize them. A pancreatic cancer-specific data screening module is used to stratify, screen, and weight the data according to the relevant genetic mechanisms of pancreatic cancer, and generate a pancreatic cancer risk feature set. The risk prediction model module is used to output pancreatic cancer risk assessment results based on the pancreatic cancer risk feature set; The risk assessment and stratification module is used to calculate the risk ratio for pancreatic cancer and classify risk levels. The intervention recommendation output module is used to output corresponding health management or medical intervention recommendations based on the risk level; The modules work together to implement the method described in claim 1.
10. A computer-readable storage medium, characterized in that, The storage medium stores multiple instructions adapted for loading by a processor to execute the steps in the pancreatic cancer risk prediction method based on pancreatic cancer-related genomic biomarkers as described in any one of claims 1 to 8.