Construction method of prediction model of prostatic cancer malignancy risk and prediction system

By collecting and analyzing the genetic variation data of the subjects and generating genetic variation fingerprints, the problem of the inability to predict the malignancy trend of prostate cancer in the prior art is solved, and the accurate prediction of the risk of malignant prostate cancer is achieved, and over-treatment of inert prostate cancer is reduced.

CN120340599APending Publication Date: 2025-07-18HANGZHOU INST FOR ADVANCED STUDY UCAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311629869.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cannot effectively predict the malignancy trend of prostate cancer, resulting in unnecessarily surgical resection of early prostate cancer subjects, resulting in excessive treatment of indolent prostate cancer subjects.

Method used

By collecting subject genetic variation data, intersecting with the SNP set of prostate cancer GWAS site in UKBB, the correlation between genetic mutations and 3βHSD1 activity was calculated, and the risk score was calculated using logistic regression model to generate genetic variation fingerprints to predict the risk of prostate cancer malignant.

Benefits of technology

It provides a method that can identify subjects at high risk of malignant progression of prostate cancer, reduces the risk of overtreatment in indolent prostate cancer subjects and improves the accuracy of early diagnosis of prostate cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340599A_ABST
    Figure CN120340599A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a prediction model for prostate cancer malignancy risk and a prediction system, and the method comprises the following steps: S1, collecting sample data of hereditary variation of a subject, including mutation allelic dose of variation; s2, taking an intersection of the sample data of the heritable variation of the subject in the step S1 and a prostatic cancer GWAS site SNP set in UKBB; s3, calculating the correlation between the genetic mutation of the intersection in S2 and the activity of 3beta HSD1; s4, calculating the risk score of the subject; and S5, randomly selecting a plurality of genetic variations for multiple times, repeating the step S4, calculating to obtain a plurality of PRSs, and calculating a genetic variation fingerprint score. According to the construction method of the prediction model for the malignant risk of the prostate cancer and the prediction system, the tissue with high 3beta HSD1 enzyme activity and the tissue with low 3beta HSD1 enzyme activity are subjected to multi-omics sequencing, and genetic variation fingerprints are established to represent the metabolic activity of the 3beta HSD1 of the prostate tissue; the biomarker can be provided for doctors as an efficient biomarker for early diagnosis of prostatic cancer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical information diagnosis and prediction, and in particular to a method for constructing a prediction model and a prediction system for the risk of prostate cancer malignancy. Background Art

[0002] The incidence of prostate cancer is high, but only about 15% of the subjects have malignant prostate cancer and need active treatment; the other 85% or so subjects have indolent prostate cancer, the disease progresses slowly, and the subjects can survive with cancer. However, in the early stages of the disease, there is currently no method to predict the malignant trend of early prostate cancer, and it is impossible to screen potential malignant prostate cancer subjects. Therefore, the guidelines recommend that early subjects should be mainly monitored clinically, and active intervention should be carried out after the disease progresses further, which objectively causes delays in the treatment of malignant prostate cancer. In clinical practice in my country, prostate surgery is often performed on subjects with early prostate cancer without distinction, resulting in overtreatment of subjects with indolent prostate cancer.

[0003] Therefore, how to provide a method for predicting the risk of prostate cancer, so as to achieve earlier and more accurate identification of subjects at high risk of developing prostate cancer and subjects at high risk of malignant progression of prostate cancer, is a technical problem that needs to be solved urgently by technicians in this field. Many studies have pointed out that the metabolic enzyme 3βHSD1 is associated with the malignant progression of advanced prostate cancer; the gene polymorphism of HSD3B1 encoding 3βHSD1 is associated with the malignant progression of advanced prostate cancer. However, the HSD3B1 (1245C) genotype associated with malignant prostate cancer appears in very low proportion in the Oriental population, which loses its value in disease prediction and diagnosis. Therefore, it is of great significance to provide a method that is suitable for Chinese patients and can indicate the activity of the metabolic enzyme 3βHSD1. Summary of the invention

[0004] The first object of the present invention is to provide a method for constructing a prediction model for the malignancy risk of prostate cancer in response to the problems in the prior art.

[0005] To this end, the above-mentioned purpose of the present invention is achieved through the following technical solutions:

[0006] A method for constructing a prediction model for the risk of prostate cancer malignancy comprises the following steps:

[0007] S1, collect sample data of genetic variations of subjects, including mutant allele dosage of the variants;

[0008] S2, taking the intersection of the sample data of the subject's genetic variation in step S1 and the set of SNPs of the prostate cancer GWAS sites in UKBB;

[0009] S3, calculate the correlation between the genetic mutations of the intersection of S2 and 3βHSD1 activity, the calculation method is:

[0010] Y = βX

[0011] The model used here is logistic regression, where Y is the 3βHSD1 activity, X is the mutant genotype, and β is the change in 3βHSD1 activity brought about by one mutant allele;

[0012] S4. Calculate the risk score of the subject. The score calculation model is

[0013] where PRS j represents the risk score of the j-th subject, and d ij represents the mutant allele dose of the i-th variant of the j-th subject, and β i represents the β of the i-th variant;

[0014] Randomly select several genetic variants multiple times, repeat step S4, calculate several PRSs. Each time, judge the 3βHSD1 activity label of the subject according to the PRS of the subject. The judgment criterion is: select the optimal PRS cut-off point to minimize the Gini value after segmentation. Then, the subjects with a PRS lower than this cut-off point are considered to have low 3βHSD1 activity, and the subjects with a PRS higher than the cut-off point are considered to have high 3βHSD1 activity. Finally, count the proportion of high 3βHSD1 labels of the subjects after multiple judgments. The proportion value is used as the genetic variant fingerprint score. When the proportion value is greater than 0.5, the subject is considered to have high 3βHSD1 activity, and when the proportion value is less than 0.5, the subject is considered to have low 3βHSD1 activity.

[0015] While adopting the above technical solution, the present invention can also adopt or combine the following technical solutions:

[0016] As a preferred technical solution of the present invention: In step S5, it includes the following steps:

[0017] S5.1. Randomly select several genetic variants and calculate the PRS of the subject,

[0018] S5.2. Calculate the Gini value of the PRS child node of the subject,

[0019] Gini(D) represents the Gini value of the child node, and p i represents the proportion of the i-th type of sample in this node, and p i′ represents the proportion of other samples except the i-th type of sample in this node,

[0020] S5.3. Select the Gini(D) with the smallest value as the optimal cut-off point of the PRS.

[0021] As a preferred technical solution of the present invention: in step S3, a genetic mutation β related to 3βHSD1 activity is selected by t-test.

[0022] As a preferred technical solution of the present invention: step S5 includes the following steps: the bootstrap algorithm is used to perform 1000 times of random sampling with replacement for both the subjects and the mutations respectively, then the subjects and the mutations are combined, the PRS of the subjects determined by the mutations in 1000 combinations of subjects and mutations is calculated, and then the proportions of the high 3βHSD1 activity tags and the low 3βHSD1 activity tags obtained for each sample in the 1000 combinations are statistically analyzed. If the proportion of the high 3βHSD1 activity tags is greater than 0.5, it is considered that the final prediction tag of the sample is high 3βHSD1 activity; on the contrary, if it is less than 0.5, it is considered that the final prediction tag of the sample is low 3βHSD1 activity. When making predictions on new data without known 3βHSD1 activity tags, 1000 mutation sets identical to the training set will be formed first, but no similar sample sets will be formed. Subsequently, the 1000 sets of sample PRS corresponding to the 1000 mutation sets in the new data are calculated, and each set of PRS is used to divide high 3βHSD1 activity and low 3βHSD1 activity using the corresponding threshold in the training set. Finally, the prediction tag is determined by statistically analyzing the proportions of the two types of tags in the 1000 sets of samples.

[0023] The second object of the present invention is to provide a prediction system based on a prostate cancer malignancy risk prediction model for the problems in the prior art.

[0024] To this end, the above object of the present invention is achieved by the following technical solutions:

[0025] A prediction system for a method of constructing a prediction model based on prostate cancer malignancy risk, characterized in that:

[0026] It includes an acquisition module configured to acquire sample data of genetic variations of a subject as input data;

[0027] A calculation module configured to input the input data into the prediction model and obtain the score of the genetic variation fingerprint as the prediction result of the prostate cancer risk.

[0028] The present invention has the following beneficial effects: A method for constructing a prediction model and a prediction system for the malignant risk of prostate cancer of the present invention, through multi-omics sequencing of tissues with high 3βHSD1 enzyme activity and tissues with low 3βHSD1 enzyme activity, establishes a specific genetic variation fingerprint to characterize the 3βHSD1 metabolic activity of prostate tissues, which can be provided to doctors as a biomarker for efficient early diagnosis of prostate cancer to judge the incidence risk and malignant progression risk of prostate cancer in corresponding subjects. The method and system for constructing a prediction model for the malignant risk of prostate cancer of the present invention, the genetic variation fingerprint can achieve: using the blood tissue of the subject to measure the variation of specific gene genetic loci, calculating the corresponding score of the metabolic enzyme 3βHSD1 activity, and using it to predict the trend of malignant progression of the disease. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 The 3βHSD1 activity of prostate tissue is related to tumor incidence and malignant tendency. Among them, A is the metabolic pathway of androgens in prostate tissue; B is the individual heterogeneity of 3βHSD1 enzyme activity in prostate tissue. On the Y-axis, the ratio of pro-carcinogenic androgens / (DHEA + pro-carcinogenic androgens) is calculated to evaluate the 3βHSD1 enzyme activity of the prostate tissue of the subject. On the X-axis, the tissue metabolism results of 239 subjects. C shows that the 3βHSD1 activity of prostate tissue in metastatic cancer subjects is high; D shows that subjects with high 3βHSD1 activity of prostate tissue are more likely to develop tolerance to castration therapy; E shows that the 3βHSD1 activity of prostate tissue in subjects taking finasteride is low. Previous clinical trials have shown that finasteride can reduce the incidence risk of prostate cancer;

[0030] Figure 2 The genetic variation fingerprint for characterizing the 3βHSD1 activity of prostate tissue. Among them, A, 31 needle biopsy tissues with high 3βHSD1 activity and 29 needle biopsy tissues with low enzyme activity are selected for multi-omics sequencing; B, the flow chart for establishing the genetic variation fingerprint for characterizing the 3βHSD1 activity of prostate tissue; C, the distribution of the genetic variation fingerprint for characterizing the 3βHSD1 activity of prostate tissue in high- and low-metabolism patients; D, using the Chinese Prostate Cancer Genome and Epigenome Atlas Database CPGEA to verify the discrimination of the genetic variation fingerprint; E is using the functional genes corresponding to the genetic variation fingerprint to verify the patient clustering effect in the TCGA database; F is using the functional genes corresponding to the genetic variation fingerprint to verify the patient clustering effect in the CPGEA database. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present invention will be further described in detail with reference to the accompanying drawings and specific embodiments.

[0032] A method for constructing a prediction model for the malignant risk of prostate cancer of the present invention includes the following steps:

[0033] S1. Collect sample data of genetic variations of the subjects, including the mutant allele dosage of the variations;

[0034] S2. Take the intersection of the sample data of the genetic variations of the subjects in step S1 and the set of SNPs at the prostate cancer GWAS loci in UKBB;

[0035] S3. Calculate the correlation between the genetic mutations in the intersection of S2 and the 3βHSD1 activity. The calculation method is:

[0036] Y = βX

[0037] The model used here is logistic regression, where Y is the 3βHSD1 activity, X is the mutant genotype, and β is the change in 3βHSD1 activity brought about by one mutant allele;

[0038] S4. Calculate the risk score of the subject. The score calculation model is

[0039] where PRS j represents the risk score of the j-th subject, d ij represents the mutant allele dosage of the i-th variation of the j-th subject, and β i represents the β of the i-th variation;

[0040] S5. Randomly select several genetic variations multiple times, repeat step S4, calculate a number of PRSs, and each time determine the 3βHSD1 activity label of the subject according to the PRS of the subject. The determination criterion is: select the optimal PRS cut-off point to minimize the Gini value after segmentation. Subjects determined to be less than this PRS cut-off point are considered to have low 3βHSD1 activity, and subjects higher than the PRS cut-off point are considered to have high 3βHSD1 activity. Finally, count the proportion of subjects with high 3βHSD1 labels after multiple determinations, and this proportion value is used as the genetic variation fingerprint score. When the proportion value is greater than 0.5, the subject is considered to have high 3βHSD1 activity, and when the proportion value is less than 0.5, the subject is considered to have low 3βHSD1 activity.

[0041] Among them, UKBB is the abbreviation of the UK Biobank, a large biomedical database and research resource that contains genetic and health information from 500,000 UK participants.

[0042] GWAS is genome-wide association study, that is, to find out the existing sequence variations SNPs in the entire human genome and the SNPs related to diseases. It is a method that uses statistical methods, namely regression analysis, to solve genetic problems and is also an important part of quantitative genetics research. In step S5, the following steps are included:

[0043] S5.1. Randomly select a number of genetic variations and calculate the PRS of the subject

[0044] S5.2. Calculate the Gini value of the PRS child node of the subject

[0045] Gini(D) represents the Gini value of the child node, and p i represents the proportion of the i-th type of sample in this node, and p i′ represents the proportion of other samples except the i-th type of sample in this node

[0046] S5.3. Select the Gini(D) with the smallest value as the optimal splitting point of the PRS

[0047] In step S3, genetic mutation β related to 3βHSD1 activity is selected through t-test

[0048] Step S5 includes the following steps: The bootstrap algorithm is used to perform 1000 times of random sampling with replacement on both the subjects and the variations respectively, and then the subjects and the variations are combined. Calculate the PRS of the subjects determined by the variations in 1000 combinations of subjects and variations. Subsequently, count the proportions of high 3βHSD1 activity labels and low 3βHSD1 activity labels obtained for each sample in the 1000 combinations. If the proportion of high 3βHSD1 activity labels is greater than 0.5, it is considered that the final prediction label of this sample is high 3βHSD1 activity. Conversely, if it is less than 0.5, it is considered that the final prediction label of this sample is low 3βHSD1 activity. When making predictions on new data without known 3βHSD1 activity labels, 1000 variation sets identical to the training set will be formed first, but no similar sample sets will be formed. Subsequently, calculate the 1000 groups of sample PRS corresponding to the 1000 variation sets in the new data. Each group of PRS is used to divide high 3βHSD1 activity and low 3βHSD1 activity using the corresponding threshold in the training set. Finally, determine the prediction label by counting the proportions of the two types of labels in the 1000 groups of samples

[0049] A prediction system for a method of constructing a prediction model based on the malignant risk of prostate cancer, including an acquisition module configured to acquire sample data of genetic variations of a subject as input data; a calculation module configured to input the input data into the prediction model and obtain the score of the genetic variation fingerprint as the prediction result of the prostate cancer risk

[0050] A method of constructing a prediction model for the malignant risk of prostate cancer according to the present invention establishes a combination of genetic variations that can reflect the metabolic activity of 3βHSD1, thereby providing a reference for doctors to diagnose prostate cancer risk prediction

[0051] A method for constructing a prediction model for the malignant risk of prostate cancer

[0052] The data used in the generation of the genetic variation fingerprint characterizing the 3βHSD1 activity of prostate tissue is the genetic mutation data of the subjects, which is derived from UKBB and the subject cohort (training set cohort).

[0053] First, 4,806,466 high-confidence genetic mutations were obtained from a cohort of 47 subjects in the training set cohort, and then overlapped with the genetic mutations related to the genetic risk of prostate cancer in the UKBB cohort to obtain 4,425 genetic mutations related to the risk of prostate cancer that exist in both the 47-subject cohort and UKBB. Subsequently, the correlation between these genetic mutations and 3βHSD1 activity was calculated. The calculation method is:

[0054] Y = βX

[0055] Where Y is the 3βHSD1 activity, X is the mutant genotype, and the regression coefficient β is measured by the t-test to determine the correlation between the genetic mutation and the 3βHSD1 activity. In this way, 28 genetic mutations significantly related to the 3βHSD1 activity were finally identified, thus forming a genetic variation fingerprint characterizing the 3βHSD1 activity of prostate tissue.

[0056] Calculate the score of the individual genetic variation fingerprint:

[0057] During the process of generating the genetic variation fingerprint characterizing the 3βHSD1 activity of prostate tissue, the influence β of each variation on the 3βHSD1 activity was obtained, and it was used as the weight of the variation. The specific scoring algorithm is:

[0058]

[0059] Where PRS j represents the risk score of the j-th subject, d ij represents the mutant allele dose of the i-th variation of the j-th subject, and β i represents the influence of the i-th variation. Thus, a risk score can be calculated for each subject, and the genetic variation fingerprint can be applied to a population cohort with unknown 3βHSD1 activity.

[0060] In the present invention, during the process of generating the genetic variant fingerprint representing the activity of 3βHSD1 in prostate tissue, the influence β of each variant on the activity of 3βHSD1 was obtained. The prediction model first had to form and fix the various parameters in the training set subject cohort. Since there was a fuzzy interval in the PRS of all samples, that is, for samples with high 3βHSD1 activity within this interval, their PRS might be greater than that of samples with low activity, or vice versa. Therefore, in order for the prediction model to better distinguish this part of the samples, a fuzzy threshold of PRS was generated again. The specific method was as follows: The bootstrap algorithm was used to perform 1000 times of random sampling with replacement for both the subject samples and the variants respectively, and then the samples and the variants were combined. The PRS of the samples determined by the variants in 1000 combinations of samples and variants was calculated. In each combination, the PRS threshold that could best distinguish the samples with high 3βHSD1 activity and low 3βHSD1 activity was selected. Therefore, each combination would assign new 3βHSD1 activity labels to the samples in the combination. Subsequently, the proportions of the high 3βHSD1 activity labels and low 3βHSD1 activity labels obtained by each sample in 1000 combinations were statistically analyzed. If the proportion of the high 3βHSD1 activity labels was greater than 0.5, it was considered that the final prediction label of the sample was high 3βHSD1 activity; conversely, if it was less than 0.5, it was considered that the final prediction label of the sample was low 3βHSD1 activity. In the subject cohort, the sensitivity of the prediction model was 0.96 and the specificity was 0.864 at a threshold of 0.5. When making predictions on new data without known 3βHSD1 activity labels, 1000 variant sets identical to those in the training set would be formed first, but no similar sample sets would be formed. Subsequently, the 1000 sets of sample PRS corresponding to the 1000 variant sets in the new data were calculated. Each set of PRS was used to divide the samples into high 3βHSD1 activity and low 3βHSD1 activity using the corresponding threshold in the training set. Finally, the prediction labels were determined by statistically analyzing the proportions of the two types of labels of 1000 sets of samples.

[0061] In the prior art, androgens drive the occurrence and development of prostate cancer. Prostate tissue can convert testosterone from the testis and dehydroepiandrosterone (DHEA) from the adrenal gland into dihydrotestosterone (DHT), further activating the androgen receptor (AR) signaling pathway and promoting the malignant progression of the disease. The research group proposed that the stronger the ability of prostate tissue to synthesize androgens, the higher the malignant trend of prostate cancer in the corresponding subjects. Since the activity of the metabolic enzyme SRD5A in the Chinese population is low and the utilization rate of testosterone from the testis is poor, the adrenal-prostate androgen metabolic axis mediated by the metabolic enzyme 3βHSD1 may be of greater significance in Chinese subjects.

[0062] The method for predicting the risk of prostate cancer uses an in vitro tissue metabolism tracing system to treat freshly obtained prostate biopsy tissues with dehydroepiandrosterone labeled with radioactive isotopes, and evaluates the utilization of dehydroepiandrosterone in the tissues of the subjects and the activities of corresponding metabolic enzymes by tracing tissue metabolism. By comparing the tissue metabolism results of subjects in different disease stages, with different clinical responses, and under different drug treatments, it is confirmed that high activity of the metabolic enzyme 3βHSD1 in the tissue is a risk factor for prostate development and cancer malignant progression.

[0063] To obtain a diagnostic method for patient stratification with greater clinical application potential, multi-omics sequencing was performed on tissues with high 3βHSD1 enzyme activity and tissues with low 3βHSD1 enzyme activity, and a genetic variation fingerprint capable of characterizing 3βHSD1 enzyme activity was established to characterize tissue metabolic activity, thereby predicting the risk of disease malignancy. The corresponding method was verified in multiple databases.

[0064] As Figure 1 shown, the 3βHSD1 activity in prostate tissue is related to the tendency of tumor malignancy. As Figure 1 A, the metabolic pathway of androgens in prostate tissue. As Figure 1 B, the individual heterogeneity of 3βHSD1 enzyme activity in prostate tissue. Y-axis, the ratio of pro-carcinogenic androgens / (DHEA + pro-carcinogenic androgens) was calculated to evaluate the 3βHSD1 enzyme activity in the prostate tissue of the subjects. X-axis, the tissue metabolism results of 239 subjects. Each point represents the metabolism result of a biopsy strip. As Figure 1 C, high 3βHSD1 activity in the prostate tissue of metastatic cancer subjects. The biopsy strips were from newly diagnosed subjects in different disease stages. Figure 1 D, subjects with high 3βHSD1 activity in prostate tissue are more likely to develop resistance to castration therapy. The subjects received a biopsy before castration therapy. Figure 1 E, finasteride inhibits high 3βHSD1 activity in prostate tissue. The DHEA metabolism results of the prostate tissues of newly diagnosed subjects taking prostate hyperplasia treatment drugs (α-blockers and finasteride). Finasteride can reduce the incidence of prostate cancer. In this subject cohort, the overall 3βHSD1 activity in the prostate tissue of subjects taking finasteride was lower than that of subjects taking α-blockers.

[0065] As Figure 2 shown, the genetic variation fingerprint characterizing 3βHSD1 activity in prostate tissue. Figure 2Among them, as shown in Figure A, 31 cases of puncture tissues with high 3βHSD1 activity and 29 cases of puncture tissues with low enzyme activity were selected for multi-omics sequencing. As shown in Figure B, the flow chart for establishing the genetic variation fingerprint characterizing the 3βHSD1 activity in prostate tissues. As shown in Figure C, the distribution of the genetic variation fingerprint characterizing the 3βHSD1 activity in prostate tissues among high- and low-metabolism patients. As shown in Figure D, the CPGEA database was used to verify the discrimination of the genetic variation fingerprint. CPGEA is a database of multi-omics sequencing results of prostate cancer patients published by Changhai Hospital. Patients with high genetic variation fingerprint scores are more likely to experience biochemical recurrence. As shown in Figures E and F, the patient clustering effect was verified using the functional genes corresponding to the genetic variation fingerprint in the TCGA and CPGEA databases. Since most databases do not provide the sequencing results of patients' blood samples, we established a gene set related to the genetic variation fingerprint through functional genomics, as shown in Figure B, and used this gene set to cluster patients in the TCGA and CPGEA. It was also found that the gene set corresponding to the genetic variation fingerprint has the function of patient clustering.

[0066] Among them, CPGEA is the abbreviation of the Chinese Prostate Cancer Genome Atlas, which is a research project initiated by Chinese scientists aiming to map the genomic and epigenomic landscapes of the Chinese prostate cancer population and explore their relationships with disease occurrence and development. The project was launched in 2015 and has currently released multiple research results, providing important reference bases for the early diagnosis, treatment, and prognosis evaluation of prostate cancer.

[0067] The Chinese name of TCGA is The Cancer Genome Atlas Project, which is a large-scale research project initiated by the US National Cancer Institute (NCI) aiming to systematically study the genomic changes in cancer and their relationships with clinical prognosis. The goal of TCGA is to comprehensively depict the genomic variation maps of all human cancers, thereby deeply understanding the mechanisms of cancer occurrence and development.

[0068] The construction method and prediction system of a prediction model for the malignant risk of prostate cancer of the present invention provide an efficient data processing method for biomarkers for the early diagnosis of prostate cancer. According to the stronger the ability of prostate tissues to synthesize androgens, the data of the activity of the metabolic enzyme 3βHSD1 is used for doctors to judge the prostate cancer incidence risk and malignant progression risk of the corresponding subjects, and then a reference data is provided to doctors to provide data support for their diagnosis of the prostate cancer incidence risk and prostate cancer malignant progression risk of the subjects.

[0069] The above specific embodiments are used to explain the present invention, which are only the preferred embodiments of the present invention and do not limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and protection scope of the claims of the present invention fall within the protection scope of the present invention.

Claims

1. A method for constructing a prediction model for the malignant risk of prostate cancer, comprising the following steps: S1. Collect sample data of genetic variations of subjects, including the mutant allele dosage of the variations; S2. Take the intersection of the sample data of the genetic variations of the subjects in step S1 and the set of SNPs at the prostate cancer GWAS loci in UKBB; S3. Calculate the correlation between the genetic mutations in the intersection of S2 and the 3βHSD1 activity. The calculation method is: Y = βX Here, logistic regression analysis is used, where Y is the 3βHSD1 activity, X is the mutant genotype, and β is the change in 3βHSD1 activity brought about by one mutant allele; S4. Calculate the risk score of the subject, and the score calculation model is where PRS j represents the risk score of the j-th subject, d ij represents the mutant allele dose of the i-th variant of the j-th subject, and β i represents β of the i-th variant; S5. Randomly select a number of genetic variations multiple times, repeat step S4, calculate a number of PRSs. Each time, judge the 3βHSD1 activity label of the subject according to the PRS of the subject. The judgment criterion is: select the optimal PRS segmentation point to minimize the Gini value after segmentation. Subjects judged to be less than this PRS segmentation point are considered to have low 3βHSD1 activity, and subjects higher than the PRS segmentation point are considered to have high 3βHSD1 activity. Finally, count the proportion of subjects with high 3βHSD1 labels after multiple judgments, and this proportion value is used as the genetic variation fingerprint score.

2. The method for constructing a prediction model for the malignant risk of prostate cancer according to claim 1, characterized in that: Step S5 specifically includes the following steps: S5.

1. Randomly select a number of genetic variations and calculate the PRS of the subject; S5.2, calculate the Gini value of the PRS child node of the subject. Gini(D) represents the Gini value of the child node, and p i represents the proportion of the i-th class samples in this node, and p i′ represents the proportion of the other samples except the i-th class samples in this node. S5.

3. Select the Gini(D) with the smallest value as the optimal segmentation point of the PRS.

3. The method for constructing a prediction model for the malignant risk of prostate cancer according to claim 1, wherein: In step S3, the genetic mutation β related to the 3βHSD1 activity is selected through a t-test.

4. The method for constructing a prediction model for the malignant risk of prostate cancer according to claim 1, wherein: Step S5 includes the following steps: The bootstrap algorithm is used to perform 1000 times of random sampling with replacement for both the subjects and the variations respectively, and then the subjects and the variations are combined. Calculate the PRS of the subjects determined by the variations in 1000 combinations of subjects and variations. Subsequently, count the proportions of high 3βHSD1 activity labels and low 3βHSD1 activity labels obtained for each sample in the 1000 combinations. If the proportion of high 3βHSD1 activity labels is greater than 0.5, it is considered that the final prediction label of the sample is high 3βHSD1 activity. Conversely, if it is less than 0.5, it is considered that the final prediction label of the sample is low 3βHSD1 activity. When making predictions on new data without known 3βHSD1 activity labels, 1000 variation sets identical to the training set will be formed first, but no similar sample sets will be formed. Subsequently, calculate the 1000 groups of sample PRSs corresponding to the 1000 variation sets in the new data. Each group of PRSs is used to divide high 3βHSD1 activity and low 3βHSD1 activity using the corresponding threshold in the training set. Finally, determine the prediction label by counting the proportions of the two types of labels in the 1000 groups of samples.

5. The method for constructing a prediction model for the malignant risk of prostate cancer according to claim 1, wherein: When the genetic variation fingerprint score is greater than 0.5, it is considered that the subject has high 3βHSD1 activity. When the genetic variation fingerprint score is less than 0.5, it is considered that the subject has low 3βHSD1 activity.

6. A prediction system based on the method for constructing a prediction model for the malignant risk of prostate cancer according to any one of claims 1-5, characterized in that: An acquisition module, configured to acquire sample data of a subject's genetic variation as input data; A calculation module, configured to input the input data into the prediction model and obtain the score of the output genetic variation fingerprint as the prediction result of the risk of prostate cancer.