Bladder cancer prediction model training method, bladder cancer prediction method equipment and medium

By training a bladder cancer prediction model based on LASSO algorithm, and using methylation data to differentially screen samples, the processing efficiency of bladder cancer detection data and model calculation efficiency are improved, the problem of inefficient bladder cancer detection in the existing technology is solved, and rapid and accurate bladder cancer prediction is achieved.

CN120236749APending Publication Date: 2025-07-01SANSURE BIOTECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311867094.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, the processing of gene detection data for bladder cancer detection is inefficient, and medical personnel need to spend a lot of time to perform manual analysis, resulting in inefficient processing of bladder cancer prediction results.

Method used

By using the LASSO algorithm to train the bladder cancer prediction model, the methylation sample data is determined using the control methylation data and the public methylation data, the training set and test set are divided, the model is trained and tested, and the model with the best performance indicators is selected as the bladder cancer prediction model.

Benefits of technology

It improves the processing efficiency of bladder cancer detection data, reduces the processing time of gene detection data, improves the calculation efficiency of the model, and can quickly and accurately predict bladder cancer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236749A_ABST
    Figure CN120236749A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a bladder cancer prediction model training method, bladder cancer prediction method equipment and a medium, and belongs to the field of biological detection. The bladder cancer prediction model training method comprises the following steps: determining a first quantity of methylation sample data according to methylation value difference between contrast methylation data and acquired public methylation data; dividing the first quantity of methylated sample data into a training set and a test set; inputting the training set into a first number of preset models, and training the preset models based on an LASSO algorithm to obtain a first number of trained models; inputting the test set into the first number of trained models, testing the trained models, and determining a performance index of each trained model; and determining a bladder cancer prediction model according to the performance indexes. And the methylation data of the at least one gene segment is determined as methylation sample data, so that the model calculation efficiency is improved, and the efficiency of processing the bladder cancer detection data is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biological detection, and specifically to a method for training a bladder cancer prediction model, a bladder cancer prediction method, device, equipment and medium. Background Art

[0002] Bladder cancer is one of the common malignant tumor diseases in the urinary system. With the rapid development of medical diagnosis technology, it is possible to detect the blood, urine and sample tissues of the target person, and obtain the gene detection data of the target person. By analyzing the gene detection data of the target person, a bladder cancer prediction result of the target person can be obtained. When the bladder cancer prediction result is that the target person has bladder cancer, it is determined whether surgical treatment is required for the target person to avoid the bladder cancer patient missing the best treatment time.

[0003] However, in the actual bladder cancer detection scenario, medical staff perform manual analysis on gene detection data. Analyzing gene detection data takes a lot of time. Each medical staff can only obtain a small number of bladder cancer prediction results in a short time, and the processing efficiency of gene detection data is low, and the bladder cancer prediction result cannot be quickly determined. In addition, in order to improve the accuracy of the bladder cancer prediction result, it is usually necessary for a target person to undergo disease detection multiple times to obtain multiple gene detection data. For the bladder cancer detection of each target person, medical staff need to spend a lot of time performing manual analysis on the multiple gene detection data, resulting in low processing efficiency of gene detection data. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method for training a bladder cancer prediction model, a bladder cancer prediction method, device, equipment and storage medium. The method for training a bladder cancer prediction model is used to solve the problem of low efficiency in processing bladder cancer detection data.

[0005] To achieve the above purpose, in the first aspect, the present application provides a method for training a bladder cancer prediction model, and the method for training a bladder cancer prediction model includes:

[0006] Determine a first quantity of methylation sample data according to the methylation value difference between the control methylation data and the obtained public methylation data, where the methylation sample data is the methylation data of at least one gene fragment in the public methylation data, and the control methylation data is the methylation data of a user without bladder cancer;

[0007] Divide the first quantity of methylation sample data into a training set and a test set;

[0008] Input the training set into a first number of preset models, train the preset models based on the LASSO algorithm to obtain a first number of trained models, where the alpha parameters of each preset model are different, and the number of gene loci in the gene fragments recognized by each trained model is different;

[0009] Input the test set into the first number of trained models, test the trained models, and determine the performance metrics of each trained model;

[0010] Determine a bladder cancer prediction model according to the performance metrics.

[0011] In the embodiments of the present application, according to the methylation value differences between the control methylation data and the obtained public methylation data, a first number of methylation sample data are determined, including:

[0012] Obtain a first number of public methylation data in the database;

[0013] Compare the public methylation data with the control methylation data to determine the methylation value differences of all gene loci in the public methylation data;

[0014] According to the methylation value differences of all gene loci, determine target gene fragments, where the target gene fragments include at least two adjacent target gene loci, the methylation value differences of all target gene loci are the same, and the methylation value difference of the target gene loci is greater than a preset difference threshold;

[0015] Determine the methylation data corresponding to the target gene fragments as methylation sample data.

[0016] In the embodiments of the present application, input the test set into the first number of trained models, test the trained models, and determine the performance metrics of each trained model, including:

[0017] Input the test set and the real sample data set into the first number of trained models respectively, test the trained models, and determine the performance metrics of each trained model, where the real sample data set is obtained by performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users.

[0018] In the embodiments of the present application, the bladder cancer prediction model training method further includes:

[0019] Perform gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users to obtain real methylation data;

[0020] Determine the methylation value corresponding to the obtained real methylation data as the real sample data of the user.

[0021] In an embodiment of the present application, the test set is input into the first number of trained models, and the trained models are tested to determine the performance metrics of each trained model, including:

[0022] Obtain a second number of sample methylation data;

[0023] Respectively input the test set, the real sample data set, and the second number of sample methylation data into the first number of trained models, test the trained models, and determine the performance metrics of each trained model.

[0024] In an embodiment of the present application, the performance metrics include a first performance metric corresponding to the test set, a second performance metric corresponding to the real sample data set, and a third performance metric corresponding to the second number of sample methylation data;

[0025] According to the performance metrics, determine a bladder cancer prediction model, including:

[0026] Based on the first performance metric, the second performance metric, and the third performance metric of each trained model, determine the average value of the performance metrics of each trained model;

[0027] Determine the trained model with the highest average value of the performance metrics as the bladder cancer prediction model.

[0028] In an embodiment of the present application, the first number of trained models includes a first model with 1 identified gene locus, a second model with 5 identified gene loci, and a third model with 9 identified gene loci.

[0029] In a second aspect, the present application provides a bladder cancer prediction method, the bladder cancer prediction method:

[0030] Obtain the target methylation data of the user;

[0031] Input the target methylation data into the bladder cancer prediction model to obtain the bladder cancer prediction result of the user, where the bladder cancer prediction model is obtained according to the above-mentioned bladder cancer prediction model training method.

[0032] In a third aspect, the present application provides a computing device, including:

[0033] A memory configured to store instructions; and

[0034] A processor configured to call instructions from the memory and be able to implement the above-mentioned bladder cancer prediction model training method or the above-mentioned bladder cancer prediction method when executing the instructions.

[0035] Fourthly, the present application provides a machine-readable storage medium with a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned method for training a bladder cancer prediction model or the above-mentioned bladder cancer prediction method is implemented.

[0036] The present application provides a method for training a bladder cancer prediction model, which includes: determining a first quantity of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data; dividing the first quantity of methylated sample data into a training set and a test set; inputting the training set into the first quantity of preset models, and training the preset models based on the LASSO algorithm to obtain the first quantity of trained models; inputting the test set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model; and determining the bladder cancer prediction model according to the performance indicators. By training the bladder cancer prediction model, compared with medical staff analyzing bladder cancer detection data, the bladder cancer prediction model processes bladder cancer detection data, improving the efficiency of processing bladder cancer detection data. At the same time, compared with processing the methylation data of the complete gene sequence, determining the methylation data of at least one gene fragment as the methylated sample data improves the model calculation efficiency, and further improves the efficiency of processing bladder cancer detection data.

[0037] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. They are used to explain the embodiments of the present invention together with the following specific implementation, but do not limit the embodiments of the present invention. In the drawings:

[0039] Figure 1 Shows a flowchart of the method for training a bladder cancer prediction model provided by an embodiment of the present application;

[0040] Figure 2 Shows a flowchart of the bladder cancer prediction method provided by an embodiment of the present application;

[0041] Figure 3 Shows a schematic structural diagram of a computing device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] The following will combine the drawings in the embodiments of the present invention to detail the specific implementation of the embodiments of the present invention. It should be understood that the specific implementation described here is only used to illustrate and explain the embodiments of the present invention, and does not limit the embodiments of the present invention.

[0043] The components of the embodiments of the present invention that are generally described and illustrated in the accompanying drawings herein can be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0044] Hereinafter, the terms "comprising", "having" and their cognates that can be used in various embodiments of the present invention are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0045] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0046] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which the various embodiments of the present invention belong. The terms (such as those defined in a general use dictionary) will be construed to have the same meaning as the contextual meaning in the relevant technical field and will not be construed to have an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present invention.

[0047] Embodiment 1

[0048] Please refer to Figure 1 , Figure 1 which shows a flowchart of the method for training a bladder cancer prediction model provided by an embodiment of the present application. Figure 1 The method for training a bladder cancer prediction model in

[0049] S110, determining a first quantity of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data, wherein the methylated sample data is the methylation data of at least one gene fragment in the public methylation data, and the control methylation data is the methylation data of users without bladder cancer.

[0050] DNA (Deoxyribo Nucleic Acid) methylation refers to the addition of methyl groups to gene sequences. Usually, the methyl group is added to the cytosine ring of cytosine (C) to form 5-methylcytosine. Methylation data refers to data on the DNA methylation status, which can be obtained through gene sequence sequencing. The methylation data of users without bladder cancer is determined as control methylation data. At the same time, multiple publicly available methylation data are obtained, where the publicly available methylation data is methylation data obtained through public databases.

[0051] When diseases such as cancer usually occur, methylation abnormalities will appear. The publicly available methylation data usually obtained is the methylation data of the complete gene sequence. According to the methylation value differences between the control methylation data and the obtained publicly available methylation data, the gene fragments with methylation abnormalities caused by diseases are determined, and then the first quantity of methylation sample data is determined. The methylation sample data is the methylation data of at least one gene fragment in the publicly available methylation data, and the methylation sample data is the methylation data of the basic fragments with high methylation value differences relative to the control methylation data.

[0052] In the embodiments of the present application, determining the first quantity of methylation sample data according to the methylation value differences between the control methylation data and the obtained publicly available methylation data includes:

[0053] Obtaining the first quantity of publicly available methylation data in the database;

[0054] Comparing the publicly available methylation data with the control methylation data to determine the methylation value differences of all gene loci in the publicly available methylation data;

[0055] According to the methylation value differences of all gene loci, target gene fragments are determined, where the target gene fragments include at least two adjacent target gene loci, the methylation value differences of all target gene loci are the same, and the methylation value differences of the target gene loci are greater than the preset difference threshold;

[0056] Determining the methylation data corresponding to the target gene fragments as the methylation sample data.

[0057] The database includes publicly available methylation data of a large number of users. Among them, the type of the database is set according to actual needs and is not limited here. For the sake of understanding, in the embodiments of the present application, the database is the TGA (The cancer genome atlas) database and the GEO (Gene Expression Omnibus) database. Obtaining the first quantity of publicly available methylation data in the database, comparing the publicly available methylation data with the control methylation data, and determining the methylation value differences of all gene loci in the publicly available methylation data.

[0058] Generally, the methylation value characterizes the methylation level, which refers to the degree of methylation at a genomic locus. The methylation value is any decimal between 0 and 1. When the methylation value of a gene locus approaches 1, it is determined that the gene locus is highly methylated. When the methylation value of a gene locus approaches 0, it is determined that the gene locus is unmethylated. By determining the methylation value differences of all gene loci in the publicly available methylation data, the role of methylation in gene regulation and disease occurrence can be determined, and further whether the user has bladder cancer.

[0059] According to the methylation value differences of all gene loci, gene fragments with high consistency and high differences are determined as target gene fragments. The target gene fragment includes at least two adjacent target gene loci, the methylation value differences of all target gene loci are the same, and the methylation value difference of the target gene locus is greater than the preset difference threshold. When comparing the publicly available methylation data with the control methylation data, the larger the number of target gene loci, the higher the consistency of the gene fragment. The higher the methylation value difference of the target gene locus, the higher the difference of the gene fragment.

[0060] The methylation data corresponding to the target gene fragment is determined as the methylation sample data, so as to perform model training and testing through the methylation sample data with high consistency and high differences to obtain a high-performance model. At the same time, compared with performing model training and testing through the methylation data of the complete gene sequence, the methylation sample data filters out the methylation data with low consistency and low differences. The model only needs the methylation data corresponding to the target gene fragment, which improves the model calculation efficiency and further improves the data efficiency.

[0061] S120, divide the first quantity of methylation sample data into a training set and a test set.

[0062] In machine learning, the training set is a data set used to train a machine learning model, and the test set is a data set used to evaluate the performance of a machine learning model. Both the training set and the test set are data sets including samples.

[0063] Divide the first quantity of methylation sample data into a training set and a test set. Specifically, in the embodiments of the present application, the methylation sample data is divided based on a ratio of 7:3. Assuming 1070 methylation sample data are obtained, 749 methylation sample data are divided into the training set, and 321 methylation sample data are divided into the test set.

[0064] It should be understood that the user data with known diagnostic results is obtained for the publicly available methylation data, where the diagnostic results include that the user has bladder cancer and the user does not have bladder cancer. The diagnostic results are used as the labels of the methylation sample data to train and test the machine learning model through the methylation sample data including the labels.

[0065] S130, input the training set into the first number of preset models, and train the preset models based on the LASSO algorithm to obtain the first number of trained models, where the alpha parameters of each preset model are different, and the number of gene loci in the gene fragments recognized by each trained model is different.

[0066] Input the training set into the first number of preset models and train each preset model separately. The LASSO (Least Absolute Shrinkage and Selection Operator) algorithm is a regularization algorithm for linear regression and related models. The alpha parameters of each preset model are different, and the preset models with different alpha parameters are trained based on the LASSO algorithm to obtain the first number of trained models.

[0067] When training the preset model based on the LASSO algorithm, the coefficients of some features are forced to be compressed to zero through the regularization term, so as to achieve feature selection. The trained model can automatically select important features, making the trained model have high generalization ability.

[0068] S140, input the test set into the first number of trained models, test the trained models, and determine the performance indicators of each trained model.

[0069] Generally, for users with diseases, only some gene loci show obvious methylation value differences compared with users without diseases. Too many or too few gene loci recognized by the model will affect the performance of the model. Since the alpha parameters of each preset model are different, the number of gene loci in the gene fragments recognized by each trained model obtained is different.

[0070] Input the test set into the first number of trained models, test the trained models, and determine the performance indicators of each trained model. Generally, the higher the performance indicator of the trained model, the more accurately the trained model can distinguish positive samples and negative samples and output accurate bladder cancer prediction results. The lower the performance indicator of the trained model, the more difficult it is for the trained model to distinguish positive samples and negative samples.

[0071] S150, determine the bladder cancer prediction model according to the performance indicators.

[0072] According to the performance metrics, determine the post-training model with the optimal performance metrics, and determine the post-training model with the optimal performance metrics as the bladder cancer prediction model. By training, a bladder cancer prediction model is obtained. Compared with medical staff analyzing bladder cancer detection data, the bladder cancer prediction model processes bladder cancer detection data, improving the efficiency of processing bladder cancer detection data. At the same time, compared with processing methylation data of the complete gene sequence, determining the methylation data of at least one gene fragment as methylation sample data improves the model calculation efficiency, and thus improves the efficiency of processing bladder cancer detection data.

[0073] In the embodiments of the present application, the test set is input into the first number of post-training models, the post-training models are tested, and the performance metrics of each post-training model are determined, including:

[0074] The test set and the true sample data set are respectively input into the first number of post-training models, the post-training models are tested, and the performance metrics of each post-training model are determined, where the true sample data set is obtained by performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users.

[0075] The true sample data set is obtained by performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users. For example, the true sample data set can be obtained by performing gene sequence sequencing on the blood data of 115 bladder cancer patients, or can be obtained by performing gene sequence sequencing on the sample tissue data of 105 bladder cancer patients, or can be obtained by performing gene sequence sequencing on the urine data of 118 bladder cancer patients, which will not be elaborated here. For ease of understanding, in the embodiments of the present application, gene sequence sequencing is performed on urine data to obtain the true sample data set.

[0076] The test set and the true sample data set are respectively input into the first number of post-training models, the post-training models are tested, and the performance metrics of each post-training model are determined. The post-training models are tested with multiple sample data to obtain more accurate performance metrics of the post-training models.

[0077] In the embodiments of the present application, the method for training the bladder cancer prediction model further includes:

[0078] Perform gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users to obtain true methylation data;

[0079] Determine the methylation value corresponding to the obtained true methylation data as the true sample data of the user.

[0080] The sample tissue data, blood data, and urine data are all unmethylated data. Gene sequence sequencing is performed on at least one of the sample tissue data, blood data, and urine data of multiple users to obtain true methylation data.

[0081] Methylation analysis is performed on the true methylation data to determine the methylation level of the true methylation data, and the methylation value of the true methylation data is extracted. Methylation analysis software can be used to perform methylation analysis on the chip-captured fragments, which will not be elaborated here.

[0082] It should be understood that the true methylation data can be filtered first to filter out the methylation data of gene fragments that do not need to be analyzed, and the filtered true methylation data is obtained. By determining whether the user has methylation abnormalities based on the methylation value of the gene locus, it is further determined whether the user has bladder cancer. The methylation value corresponding to the obtained true methylation data is determined as the true sample data of the user.

[0083] In the embodiments of the present application, the test set is input into the first number of trained models, and the trained models are tested to determine the performance indicators of each trained model, including:

[0084] Obtain the second number of sample methylation data;

[0085] The test set, the true sample data set, and the second number of sample methylation data are respectively input into the first number of trained models, and the trained models are tested to determine the performance indicators of each trained model.

[0086] The second number of sample methylation data is obtained. The values of the first number and the second number are both set according to actual needs and are not limited here. For ease of understanding, in the embodiments of the present application, the first number is 1069, and the second number is 42, and there are no identical sample methylation data between the first number of sample methylation data and the second number of sample methylation data.

[0087] It should be understood that the sample methylation data includes the sample methylation data of users labeled as having bladder cancer and the sample methylation data of users not having bladder cancer. Through model training with the sample methylation data, the trained model can identify the methylation data of users having bladder cancer and the methylation data of users not having bladder cancer.

[0088] The test set, the real sample data set, and the methylation data of the second quantity of samples are respectively input into the first quantity of trained models to test the trained models, and the performance indicators of each trained model are determined. When evaluating the performance of the trained models, the real sample data set and the methylation data of the second quantity of samples are added to test the trained models, so as to verify whether the trained models can accurately distinguish positive samples and negative samples, and further obtain more accurate performance indicators.

[0089] In the embodiments of the present application, the performance indicators include the first performance indicator corresponding to the test set, the second performance indicator corresponding to the real sample data set, and the third performance indicator corresponding to the methylation data of the second quantity of samples;

[0090] According to the performance indicators, a bladder cancer prediction model is determined, including:

[0091] Based on the first performance indicator, the second performance indicator, and the third performance indicator of each trained model, the average value of the performance indicators of each trained model is determined;

[0092] The trained model with the highest average value of the performance indicators is determined as the bladder cancer prediction model.

[0093] Each trained model is tested using the test set to determine the first performance indicator of each trained model. Each trained model is tested using the real sample data set to determine the second performance indicator of each trained model. Each trained model is tested using the methylation data of the second quantity of samples to determine the third performance indicator of each trained model. Based on the first performance indicator, the second performance indicator, and the third performance indicator of each trained model, the average value of the performance indicators of each trained model is determined.

[0094] It should be understood that the average value of the performance indicators can be the average of the first performance indicator, the second performance indicator, and the third performance indicator, or the weighted average of the first performance indicator, the second performance indicator, and the third performance indicator, which will not be elaborated here. In this embodiment, the weights of the first performance indicator, the second performance indicator, and the third performance indicator are determined respectively. Specifically, the second performance indicator can be a high-weight performance indicator, and the weights of the first performance indicator and the third performance indicator can be low-weight performance indicators, and the average value of the performance indicators of each trained model is determined. The trained model with the highest average value of the performance indicators is determined as the bladder cancer prediction model, so as to determine the trained model with the best performance as the bladder cancer prediction model, so that the bladder cancer prediction model can output high-confidence results.

[0095] In the embodiments of the present application, the first quantity of trained models includes a first model with 1 identified gene locus, a second model with 5 identified gene loci, and a third model with 9 identified gene loci.

[0096] Specifically, when the alpha parameter of the preset model is 0.01, the preset model is trained to obtain the first model, and the gene locus identified by the first model is cg26112797. When the alpha parameter of the preset model is 0.02, the preset model is trained to obtain the second model, and the gene loci identified by the second model are cg26112797, cg06153925, cg10351284, cg10590292, and cg23359665. When the alpha parameter of the preset model is 0.1, the preset model is trained to obtain the third model, and the gene loci identified by the third model are cg26112797, cg06153925, cg10351284, cg10590292, cg23359665, cg00025044, cg06390079, cg11436362, and cg24757533. The first model, the second model, and the third model are respectively tested using the test set, the real sample data set, and the methylation data of the second quantity of samples to determine the performance metrics of each trained model.

[0097]

[0098]

[0099] Table 1 Performance Metrics of the Model

[0100] Please refer to Table 1, which shows the performance metrics of the first model, the second model, and the third model respectively. For ease of understanding, the performance metric in the embodiments of the present application is AUC (Area Under the Curve). AUC refers to the probability that a machine learning model ranks a positive sample before a negative sample when randomly selecting a positive sample and a negative sample. When AUC approaches 1, it is determined that the machine learning model can well distinguish positive samples and negative samples. When AUC approaches 0.5, it is determined that the machine learning model cannot effectively distinguish positive samples and negative samples. In this embodiment, according to the mean value of the performance metrics, the third model is determined as the bladder cancer prediction model.

[0101] The present application provides a method for training a bladder cancer prediction model. The method for training a bladder cancer prediction model includes: determining a first quantity of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data; dividing the first quantity of methylated sample data into a training set and a test set; inputting the training set into a first quantity of preset models, and training the preset models based on the LASSO algorithm to obtain a first quantity of trained models; inputting the test set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model; and determining a bladder cancer prediction model according to the performance indicators. By training the bladder cancer prediction model, compared with medical staff analyzing bladder cancer detection data, the bladder cancer prediction model processes bladder cancer detection data, improving the efficiency of processing bladder cancer detection data. At the same time, compared with processing methylation data of a complete gene sequence, determining the methylation data of at least one gene fragment as methylated sample data improves the model calculation efficiency, and thus improves the efficiency of processing bladder cancer detection data.

[0102] Embodiment 2

[0103] Please refer to Figure 2 , Figure 2 which shows a flowchart of the bladder cancer prediction method provided by the embodiment of the present application. Figure 2 The bladder cancer prediction method in

[0104] S210, obtaining the target methylation data of the user.

[0105] In an actual bladder cancer detection scenario, usually at least one of the user's sample tissue data, blood data, and urine data is obtained. For ease of understanding, in the embodiment of the present application, the user's urine data is obtained, and gene sequence sequencing is performed on the urine data to obtain the target methylation data of the user.

[0106] S220, inputting the target methylation data into the bladder cancer prediction model to obtain the bladder cancer prediction result of the user, where the bladder cancer prediction model is obtained according to the above-mentioned method for training a bladder cancer prediction model.

[0107] Input the target methylation data into the bladder cancer prediction model to obtain the bladder cancer prediction result of the user. Specifically, the bladder cancer prediction model outputs the result through f(x) = wTx + b, where x is the target methylation data, w and b are the parameters of the bladder cancer prediction model, which will not be elaborated here, and T represents the transpose of w. The output result of the bladder cancer prediction model is that the user has bladder cancer or the user does not have bladder cancer, and the output result obtained from the bladder cancer prediction model is determined as the bladder cancer prediction result of the user. Compared with medical staff analyzing the methylation data, the bladder cancer prediction model processes the methylation data, improving the efficiency of the methylation data, and thus can quickly obtain the bladder cancer prediction result of the user.

[0108] Please refer to Figure 3 , Figure 3 which shows the structural schematic diagram of the computing device provided by the embodiment of the present application. The embodiment of the present application also provides a computing device 300, including:

[0109] A memory 310 configured to store instructions; and

[0110] A processor 320 configured to call instructions from the memory 310 and be able to implement the above-mentioned bladder cancer prediction model training method or the above-mentioned bladder cancer prediction method when executing the instructions.

[0111] The memory 310 may include a non-permanent memory 310 in a computer-readable medium, a random access memory 310 (RAM) and / or a non-volatile memory in the form of, for example, a read-only memory 310 (ROM) or a flash memory (flash RAM), and the memory 310 includes at least one storage chip.

[0112] The processor 320 contains a kernel, and the kernel retrieves the corresponding program unit from the memory 310. One or more kernels can be set, and the problem of low efficiency in processing bladder cancer detection data can be solved by adjusting the kernel parameters.

[0113] When implementing the above-mentioned bladder cancer prediction model training method, the processor 320 is configured to:

[0114] Determine the first quantity of methylation sample data according to the methylation value difference between the control methylation data and the obtained public methylation data, where the methylation sample data is the methylation data of at least one gene fragment in the public methylation data, and the control methylation data is the methylation data of a user without bladder cancer;

[0115] Divide the first quantity of methylation sample data into a training set and a test set;

[0116] Input the training set into a first number of preset models, and train the preset models based on the LASSO algorithm to obtain a first number of trained models. Among them, the alpha parameters of each preset model are different, and the number of gene loci in the gene fragments identified by each trained model is different;

[0117] Input the test set into the first number of trained models, test the trained models, and determine the performance metrics of each trained model;

[0118] Determine the bladder cancer prediction model according to the performance metrics.

[0119] In an embodiment of the present application, the processor 320 is configured to determine a first number of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data, including: the processor 320 is configured to:

[0120] Obtain a first number of public methylation data in the database;

[0121] Compare the public methylation data with the control methylation data to determine the methylation value difference of all gene loci in the public methylation data;

[0122] Determine the target gene fragment according to the methylation value difference of all gene loci. Among them, the target gene fragment includes at least two adjacent target gene loci, the methylation value differences of all target gene loci are the same, and the methylation value difference of the target gene loci is greater than the preset difference threshold;

[0123] Determine the methylation data corresponding to the target gene fragment as the methylated sample data.

[0124] In an embodiment of the present application, the processor 320 is configured to input the test set into the first number of trained models, test the trained models, and determine the performance metrics of each trained model, including: the processor 320 is configured to:

[0125] Input the test set and the true sample data set into the first number of trained models respectively, test the trained models, and determine the performance metrics of each trained model. Among them, the true sample data set is obtained by performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users.

[0126] In an embodiment of the present application, the processor 320 is further configured to:

[0127] Perform gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users to obtain true methylation data;

[0128] Determine the methylation value corresponding to the obtained real methylation data as the real sample data of the user.

[0129] In an embodiment of the present application, the processor 320 is configured to input a test set into a first number of trained models, test the trained models, and determine the performance metrics of each trained model, including: the processor 320 is configured to:

[0130] Obtain a second number of sample methylation data;

[0131] Input the test set, the real sample data set, and the second number of sample methylation data into the first number of trained models respectively, test the trained models, and determine the performance metrics of each trained model.

[0132] In an embodiment of the present application, the performance metrics include a first performance metric corresponding to the test set, a second performance metric corresponding to the real sample data set, and a third performance metric corresponding to the second number of sample methylation data;

[0133] The processor 320 is configured to determine a bladder cancer prediction model according to the performance metrics, including: the processor 320 is configured to:

[0134] Based on the first performance metric, the second performance metric, and the third performance metric of each trained model, determine the average value of the performance metrics of each trained model;

[0135] Determine the trained model with the highest average value of the performance metrics as the bladder cancer prediction model.

[0136] In an embodiment of the present application, the first number of trained models includes a first model with 1 gene locus identified, a second model with 5 gene loci identified, and a third model with 9 gene loci identified.

[0137] When implementing the above-mentioned bladder cancer prediction method, the processor 320 is configured to:

[0138] Obtain the target methylation data of the user;

[0139] Input the target methylation data into the bladder cancer prediction model to obtain the bladder cancer prediction result of the user, where the bladder cancer prediction model is obtained according to the above-mentioned bladder cancer prediction model training method.

[0140] An embodiment of the present application further provides a machine-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned bladder cancer prediction model training method or the above-mentioned bladder cancer prediction method is implemented.

[0141] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0142] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0145] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0146] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory. The memory is an example of computer-readable media.

[0147] A machine-readable storage medium includes permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0148] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0149] The above are only embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for training a bladder cancer prediction model, characterized in that, The method for training the bladder cancer prediction model includes: Determining a first quantity of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data, where the methylated sample data is the methylation data of at least one gene fragment in the public methylation data, and the control methylation data is the methylation data of users without bladder cancer; Dividing the first quantity of the methylated sample data into a training set and a test set; Inputting the training set into a first quantity of preset models, and training the preset models based on the Least Absolute Shrinkage and Selection Operator (LASSO) algorithm to obtain a first quantity of trained models, where the alpha parameter of each preset model is different, and the number of gene loci in the gene fragments identified by each trained model is different; Inputting the test set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model; Determining the bladder cancer prediction model according to the performance indicators.

2. The method for training a bladder cancer prediction model according to claim 1, wherein The determining a first quantity of methylated sample data according to the methylation value difference between the control methylation data and the obtained public methylation data includes: Obtaining a first quantity of public methylation data in the database; Comparing the public methylation data with the control methylation data to determine the methylation value difference of all gene loci in the public methylation data; Determining target gene fragments according to the methylation value difference of all gene loci, where the target gene fragments include at least two adjacent target gene loci, the methylation value difference of all the target gene loci is the same, and the methylation value difference of the target gene loci is greater than a preset difference threshold; Determining the methylation data corresponding to the target gene fragments as the methylated sample data.

3. The method for training a bladder cancer prediction model according to claim 1, wherein The inputting the test set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model includes: Respectively inputting the test set and the true sample data set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model, where the true sample data set is obtained by performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users.

4. The method for training a bladder cancer prediction model according to claim 3, wherein The method for training the bladder cancer prediction model further includes: Performing gene sequence sequencing on at least one of the sample tissue data, blood data, and urine data of multiple users to obtain true methylation data; Determining the methylation value corresponding to the obtained true methylation data as the true sample data of the user.

5. The method for training a bladder cancer prediction model according to claim 3, wherein The inputting the test set into the first quantity of trained models, testing the trained models, and determining the performance indicators of each trained model includes: Obtaining a second quantity of sample methylation data; Input the test set, the true sample data set, and the methylation data of the second number of samples into the first number of trained models respectively to test the trained models and determine the performance metrics of each trained model.

6. The method for training a bladder cancer prediction model according to claim 5, wherein The performance metrics include a first performance metric corresponding to the test set, a second performance metric corresponding to the true sample data set, and a third performance metric corresponding to the methylation data of the second number of samples. Determining the bladder cancer prediction model according to the performance metrics includes: Based on the first performance metric, the second performance metric, and the third performance metric of each trained model, determine the average value of the performance metrics of each trained model. Determine the trained model with the highest average value of the performance metrics as the bladder cancer prediction model.

7. The method for training a bladder cancer prediction model according to claim 1, wherein The first number of trained models include a first model with 1 gene locus identified, a second model with 5 gene loci identified, and a third model with 9 gene loci identified.

8. A method for predicting bladder cancer, characterized in that, The bladder cancer prediction method: Obtain the target methylation data of the user. Input the target methylation data into the bladder cancer prediction model to obtain the bladder cancer prediction result of the user, where the bladder cancer prediction model is trained according to the bladder cancer prediction model training method described in any one of claims 1 to 7.

9. A computing device, characterized in that, Includes: A memory configured to store instructions; And A processor configured to call the instructions from the memory and be able to implement the bladder cancer prediction model training method described in any one of claims 1 to 7, or the bladder cancer prediction method described in claim 8 when executing the instructions.

10. A machine-readable storage medium, characterized in that, A computer program is stored on the machine-readable storage medium, and when the computer program is executed by the processor, it implements the bladder cancer prediction model training method described in any one of claims 1 to 7, or the bladder cancer prediction method described in claim 8.