Training Method and System for Cancer Risk Prediction Model Based on Co-Optimization
By obtaining cfDNA sequencing data and methylation characteristics of multiple cancer types, and using classification algorithms and type correlation parameter optimization training, a unified cancer risk prediction model is formed, which solves the shortcomings of early cancer screening and risk assessment in the existing technology, and achieves efficient and accurate risk prediction for different cancer types.
Patent Information
- Application Number
- CN202510505876.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In the prior art, early cancer screening and risk assessment methods have problems with insufficient detection sensitivity, limited scope of application or strong operational invasiveness, and the existing prediction models based on cfDNA methylation information fail to effectively utilize the correlation between different cancer types, resulting in insufficient generalization ability and cross-type prediction accuracy.
By obtaining cfDNA sequencing data and methylation characteristics of multiple cancer types, a classification algorithm is used to determine the data feature set of each cancer type, and a prediction model is obtained based on the type correlation parameter optimization training, and a joint optimization training is carried out in combination with the correlation parameters between different cancer types to form a unified cancer risk prediction model.
The model's ability to distinguish and generalize different cancer types has been improved, and the accuracy and adaptability of predicting cancer risks for any cancer type has been enhanced, providing scientific and reliable support for individualized cancer screening and early intervention.
Smart Images

Figure CN120108732B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a training method and system for a cancer risk prediction model based on co-optimization. Background Art
[0002] In the prior art, early screening and risk assessment of cancer usually rely on imaging examinations, blood biomarker detections, or histopathological analyses. However, these methods often have problems such as insufficient detection sensitivity, limited applicable scope, or strong operation invasiveness. With the development of high-throughput sequencing technology, cfDNA methylation information has gradually been applied to the auxiliary detection of cancer. However, existing methods mostly focus on the methylation feature modeling of a single cancer type, ignoring the potential correlations between different cancer types, resulting in deficiencies in the generalization ability and cross-type prediction accuracy of the trained prediction model, and it is difficult to achieve unified and efficient prediction of multiple cancers. It can be seen that there are defects in the prior art and urgent solutions are needed. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a training method and system for a cancer risk prediction model based on co-optimization, which can improve the discrimination ability and generalization ability of the model for different cancer types, enhance the prediction accuracy and adaptability of the cancer risk of any cancer type, and provide a more scientific and reliable model support for individualized cancer screening and early intervention.
[0004] To solve the above technical problem, in the first aspect of the present invention, a training method for a cancer risk prediction model based on co-optimization is disclosed, and the method includes:
[0005] Obtain cfDNA sequencing data and corresponding methylation features of multiple patients of multiple cancer types;
[0006] Based on a classification algorithm, determine a data feature set corresponding to each cancer type according to the cfDNA sequencing data and the corresponding methylation features;
[0007] Determine the type correlation degree parameter between any two cancer types;
[0008] Optimize and train according to the type correlation degree parameter and the data feature set corresponding to each cancer type to obtain a prediction model for predicting the cancer risk corresponding to any cancer type.
[0009] As an optional implementation manner, in the first aspect of the present invention, the methylation feature includes the methylation site feature of at least one region of the cfDNA sequencing data; the methylation site feature includes at least one of methylation rate, methylation site continuous feature, methylation site change feature, and methylation site map feature.
[0010] As an optional implementation manner, in the first aspect of the present invention, the cancer types include at least one of liver cancer, lung cancer, gastric cancer, colorectal cancer, breast cancer, and nasopharyngeal cancer.
[0011] As an optional implementation manner, in the first aspect of the present invention, based on the classification algorithm, according to the cfDNA sequencing data and the corresponding methylation characteristics, determine the data feature set corresponding to each of the cancer types, including:
[0012] For each of the patients, when the cancer type of the patient is a single cancer type, classify the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set corresponding to the corresponding cancer type;
[0013] When the cancer type of the patient is a multi-cancer type, based on the patient's medical history and the cfDNA sequencing data and the corresponding methylation characteristics, determine the closest cancer type corresponding to the patient, and classify the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set of the closest cancer type.
[0014] As an optional implementation manner, in the first aspect of the present invention, based on the patient's medical history and the cfDNA sequencing data and the corresponding methylation characteristics, determine the closest cancer type corresponding to the patient, including:
[0015] Determine the reference feature data corresponding to each cancer type included in the cancer types of the patient;
[0016] Calculate the similarity between the cfDNA sequencing data and the corresponding methylation characteristics of the patient and the reference feature data corresponding to each cancer type, and obtain the similarity parameter corresponding to each cancer type;
[0017] Calculate the confirmed diagnosis duration corresponding to each cancer type in the patient's medical history, and calculate the duration weight proportional to the confirmed diagnosis duration corresponding to each cancer type;
[0018] Calculate the product of the similarity parameter corresponding to each cancer type and the duration weight;
[0019] Determine the cancer type with the highest product as the closest cancer type corresponding to the patient.
[0020] As an optional implementation manner, in the first aspect of the present invention, determining the type correlation degree parameter between any two of the cancer types includes:
[0021] Obtain the historical patient data corresponding to each of the cancer types;
[0022] For any two of the cancer types, determine a type correlation parameter between the two cancer types according to the historical patient data of the two cancer types and the data feature set.
[0023] As an optional implementation manner, in the first aspect of the present invention, the determining a type correlation parameter between the two cancer types according to the historical patient data of the two cancer types and the data feature set includes:
[0024] For any two of the cancer types, calculate the proportion of the number of identical patients between the historical patient data of the two cancer types;
[0025] Calculate the data similarity between the data feature sets corresponding to the two cancer types;
[0026] Calculate the weighted sum average of the proportion of the number of identical patients and the data similarity to obtain the type correlation parameter between the two cancer types.
[0027] As an optional implementation manner, in the first aspect of the present invention, the optimizing and training a prediction model for predicting the cancer risk corresponding to any one of the cancer types according to the type correlation parameter and the data feature set corresponding to each cancer type includes:
[0028] Use the data feature set corresponding to each cancer type as training data labeled with the corresponding cancer type, and input it into the basic model corresponding to each cancer type for training. During the training process, jointly optimize and train the basic models corresponding to all cancer types based on a common loss function until convergence to obtain a prediction model corresponding to each cancer type; the common loss function is the weighted sum average of the loss function values of each basic model, where the weight of the loss function value corresponding to each basic model is proportional to the average value of all the type correlation parameters corresponding to the cancer type corresponding to the basic model;
[0029] Determine all the prediction models corresponding to the cancer types as a common prediction model, and the common prediction model is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any one of the cancer types.
[0030] A second aspect of the embodiments of the present invention discloses a training system for a cancer risk prediction model based on joint optimization, and the system includes:
[0031] An acquisition module, configured to acquire cfDNA sequencing data and corresponding methylation characteristics of multiple patients of multiple cancer types;
[0032] A first determination module, configured to determine a data feature set corresponding to each of the cancer types based on a classification algorithm according to the cfDNA sequencing data and the corresponding methylation features;
[0033] A second determination module, configured to determine a type correlation parameter between any two of the cancer types;
[0034] A training module, configured to optimize and train a prediction model for predicting the cancer risk corresponding to any one of the cancer types according to the type correlation parameter and the data feature set corresponding to each of the cancer types.
[0035] As an optional implementation manner, in the second aspect of the present invention, the methylation features include methylation site features of at least one region of the cfDNA sequencing data; the methylation site features include at least one of a methylation rate, a methylation site continuous feature, a methylation site change feature, and a methylation site map feature.
[0036] As an optional implementation manner, in the second aspect of the present invention, the cancer types include at least one of liver cancer, lung cancer, gastric cancer, colorectal cancer, breast cancer, and nasopharyngeal cancer.
[0037] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the first determination module determines the data feature set corresponding to each of the cancer types based on a classification algorithm according to the cfDNA sequencing data and the corresponding methylation features includes:
[0038] For each patient, when the cancer type of the patient is a single cancer type, the cfDNA sequencing data and the corresponding methylation features of the patient are classified into the data feature set corresponding to the corresponding cancer type;
[0039] When the cancer type of the patient is a multi-cancer type, based on the patient's medical history and the cfDNA sequencing data and the corresponding methylation features, determine the closest cancer type corresponding to the patient, and classify the cfDNA sequencing data and the corresponding methylation features of the patient into the data feature set of the closest cancer type.
[0040] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the first determination module determines the closest cancer type corresponding to the patient based on the patient's medical history and the cfDNA sequencing data and the corresponding methylation features includes:
[0041] Determine the reference feature data corresponding to each cancer type included in the cancer type of the patient;
[0042] Calculate the similarity between the cfDNA sequencing data and the corresponding methylation characteristics of the patient and the reference characteristic data corresponding to each cancer type, and obtain the similarity parameter corresponding to each cancer type;
[0043] Calculate the duration of disease diagnosis corresponding to each cancer type in the patient's medical history, and calculate the duration weight proportional to the duration of disease diagnosis corresponding to each cancer type;
[0044] Calculate the product of the similarity parameter corresponding to each cancer type and the duration weight;
[0045] Determine the cancer type with the highest product as the closest cancer type corresponding to the patient.
[0046] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the second determination module determines the type association degree parameter between any two of the cancer types includes:
[0047] Obtain the historical patient data corresponding to each of the cancer types;
[0048] For any two of the cancer types, determine the type association degree parameter between the two cancer types according to the historical patient data of the two cancer types and the data feature set.
[0049] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the second determination module determines the type association degree parameter between the two cancer types according to the historical patient data of the two cancer types and the data feature set includes:
[0050] For any two of the cancer types, calculate the proportion of the number of identical patients between the historical patient data of the two cancer types;
[0051] Calculate the data similarity between the data feature sets corresponding to the two cancer types;
[0052] Calculate the weighted sum average of the proportion of the number of identical patients and the data similarity to obtain the type association degree parameter between the two cancer types.
[0053] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the training module optimizes and trains a prediction model for predicting the cancer risk corresponding to any one of the cancer types according to the type association degree parameter and the data feature set corresponding to each cancer type includes:
[0054] Taking the data feature set corresponding to each of the cancer types as training data labeled with the corresponding cancer type, and inputting it into the base model corresponding to each of the cancer types for training. During the training process, based on a common loss function, the base models corresponding to all cancer types are jointly optimized and trained until convergence to obtain a prediction model corresponding to each of the cancer types; the common loss function is the weighted sum average of the loss function values of each of the base models, where the weight of the loss function value corresponding to each of the base models is proportional to the average value of all the type correlation parameters corresponding to the cancer type corresponding to the base model.
[0055] Determining the prediction models corresponding to all the cancer types as a common prediction model, and the common prediction model is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any one of the cancer types.
[0056] The third aspect of the present invention discloses another training system for a cancer risk prediction model based on joint optimization, and the system includes:
[0057] A memory storing executable program code;
[0058] A processor coupled to the memory;
[0059] The processor calls the executable program code stored in the memory and executes some or all of the steps in the training method of the cancer risk prediction model based on joint optimization disclosed in the first aspect of the present invention.
[0060] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute some or all of the steps in the training method of the cancer risk prediction model based on joint optimization disclosed in the first aspect of the present invention when being called.
[0061] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0062] The present invention obtains the sequencing data and corresponding methylation characteristics of patients with multiple cancer types, extracts the characteristic patterns of various cancers through a classification algorithm, and optimizes and trains in combination with the correlation parameters between different cancer types to obtain a prediction model for predicting the cancer risk corresponding to any one of the cancer types, so as to improve the discrimination ability and generalization ability of the model for different cancer types, enhance the prediction accuracy and adaptability of the cancer risk of any one cancer type, and provide more scientific and reliable model support for individualized cancer screening and early intervention. Description of the Drawings
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a schematic flowchart of a method for training a cancer risk prediction model based on co-optimization disclosed in an embodiment of the present invention.
[0065] Figure 2 It is a schematic structural diagram of a system for training a cancer risk prediction model based on co-optimization disclosed in an embodiment of the present invention.
[0066] Figure 3 It is a schematic structural diagram of another system for training a cancer risk prediction model based on co-optimization disclosed in an embodiment of the present invention. Detailed implementation manners
[0067] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0068] The terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or equipment.
[0069] Referring to "embodiment" in this article means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0070] The present invention discloses a training method and system for a cancer risk prediction model based on co-optimization, which obtains sequencing data and corresponding methylation features of patients with multiple cancer types, extracts feature patterns of various cancers through a classification algorithm, and optimizes and trains in combination with the correlation degree parameters between different cancer types to obtain a prediction model for predicting the cancer risk corresponding to any one of the cancer types, so as to improve the discrimination ability and generalization ability of the model for different cancer types, enhance the prediction accuracy and adaptability of the cancer risk of any cancer type, and provide a more scientific and reliable model support for individualized cancer screening and early intervention. The following will be described in detail respectively.
[0071] Embodiment 1
[0072] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a training method for a cancer risk prediction model based on co-optimization disclosed in an embodiment of the present invention. Among them, Figure 1 the described training method for the cancer risk prediction model based on co-optimization can be applied to a data processing system / data processing device / data processing server (wherein, the server includes a local processing server or a cloud processing server). As Figure 1 shown, the training method for the cancer risk prediction model based on co-optimization may include the following operations:
[0073] 101. Obtain cfDNA sequencing data and corresponding methylation features of multiple patients with multiple cancer types.
[0074] 102. Based on a classification algorithm, determine a data feature set corresponding to each cancer type according to the cfDNA sequencing data and the corresponding methylation features.
[0075] 103. Determine the type correlation degree parameter between any two cancer types.
[0076] 104. Optimize and train according to the type correlation degree parameter and the data feature set corresponding to each cancer type to obtain a prediction model for predicting the cancer risk corresponding to any cancer type.
[0077] It can be seen that the above-mentioned invention embodiments obtain sequencing data and corresponding methylation features of patients with multiple cancer types, extract feature patterns of various cancers through a classification algorithm, and optimize and train in combination with the correlation degree parameters between different cancer types to obtain a prediction model for predicting the cancer risk corresponding to any one of the cancer types, so as to improve the discrimination ability and generalization ability of the model for different cancer types, enhance the prediction accuracy and adaptability of the cancer risk of any cancer type, and provide a more scientific and reliable model support for individualized cancer screening and early intervention.
[0078] As an optional embodiment, in the above steps, the methylation features include methylation site features of at least one region of cfDNA sequencing data; the methylation site features include at least one of methylation rate, continuous methylation site features, changing methylation site features, and methylation site map features.
[0079] It can be seen that through the above optional embodiments, the content of the methylation features is defined to comprehensively characterize the methylation-related features of patients, facilitating subsequent type association calculations and model training, assisting in improving the discrimination ability and generalization ability of the model for different cancer types, enhancing the prediction accuracy and adaptability of the cancer risk of any cancer type, and providing a more scientific and reliable model support for individualized cancer screening and early intervention.
[0080] As an optional embodiment, in the above steps, the cancer types include at least one of liver cancer, lung cancer, gastric cancer, colorectal cancer, breast cancer, and nasopharyngeal cancer.
[0081] It can be seen that through the above optional embodiments, the types of cancer types are defined, facilitating subsequent type association calculations and model training, assisting in improving the discrimination ability and generalization ability of the model for different cancer types, enhancing the prediction accuracy and adaptability of the cancer risk of any cancer type, and providing a more scientific and reliable model support for individualized cancer screening and early intervention.
[0082] As an optional embodiment, in the above steps, based on a classification algorithm, according to the cfDNA sequencing data and the corresponding methylation features, a data feature set corresponding to each cancer type is determined, including:
[0083] For each patient, when the cancer type of the patient is a single cancer type, the cfDNA sequencing data and the corresponding methylation features of the patient are classified into the data feature set of the corresponding cancer type;
[0084] When the cancer type of the patient is a multi-cancer type, based on the patient's medical history, cfDNA sequencing data, and the corresponding methylation features, the closest cancer type corresponding to the patient is determined, and the cfDNA sequencing data and the corresponding methylation features of the patient are classified into the data feature set of the closest cancer type.
[0085] It can be seen that through the above optional embodiments, the cancer type of each patient is judged. When it is a single cancer type, the cfDNA sequencing data and methylation characteristics of the patient are directly classified into the data feature set corresponding to the type. When it is a multi-cancer type, the patient's medical history and methylation characteristics are combined to determine the closest cancer type, and the relevant data are classified into the data feature set of this type. Therefore, the accuracy and biological rationality of classification are improved when constructing the data feature set, which helps to enhance the adaptability and generalization ability of the subsequent prediction model to complex clinical situations, and improve the scientificity and stability of multi-cancer type risk prediction.
[0086] As an optional embodiment, in the above steps, based on the patient's medical history, cfDNA sequencing data, and corresponding methylation characteristics, determining the closest cancer type corresponding to the patient includes:
[0087] Determining the reference feature data corresponding to each cancer type included in the cancer type of the patient;
[0088] Calculating the similarity between the patient's cfDNA sequencing data and corresponding methylation characteristics and the reference feature data corresponding to each cancer type to obtain the similarity parameter corresponding to each cancer type;
[0089] Calculating the confirmed diagnosis duration corresponding to each cancer type in the patient's medical history, and calculating the duration weight proportional to the confirmed diagnosis duration corresponding to each cancer type;
[0090] Calculating the product of the similarity parameter and the duration weight corresponding to each cancer type;
[0091] Determining the cancer type with the highest product as the closest cancer type corresponding to the patient.
[0092] It can be seen that through the above optional embodiments, for multi-cancer type patients, by obtaining the reference feature data of their corresponding cancer types, calculating the similarity between their cfDNA sequencing data and methylation characteristics and each cancer type respectively, and generating a duration weight by combining the confirmed diagnosis duration information in the medical history, further determining the closest cancer type based on the product of the similarity parameter and the duration weight, so as to achieve the accurate classification of multi-cancer type patients, effectively improve the accuracy of subsequent data feature set construction, and enhance the recognition ability and classification effect of the cancer risk prediction model for multi-source heterogeneous samples.
[0093] As an optional embodiment, in the above steps, determining the type correlation degree parameter between any two cancer types includes:
[0094] Obtaining the historical patient data corresponding to each cancer type;
[0095] For any two cancer types, determine the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types.
[0096] It can be seen that through the above optional embodiments, based on the historical patient data of each cancer type and combined with its corresponding data feature set, calculate the type correlation parameter between any two cancer types, so as to accurately reflect the similarity between different cancer types at the methylation feature level, effectively improve the discrimination ability of cancer classification, provide an accurate basis for introducing type correlation information into the subsequent prediction model, thereby optimizing the model structure and improving the accuracy and generalization ability of cancer risk prediction.
[0097] As an optional embodiment, in the above steps, determining the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types includes:
[0098] For any two cancer types, calculate the proportion of the number of identical patients between the historical patient data of the two cancer types;
[0099] Calculate the data similarity between the corresponding data feature sets of the two cancer types;
[0100] Calculate the weighted sum average of the proportion of the number of identical patients and the data similarity to obtain the type correlation parameter between the two cancer types.
[0101] It can be seen that through the above optional embodiments, based on the historical patient data between any two cancer types, calculate the proportion of the number of identical patients, and combined with the similarity between the corresponding data feature sets, determine the type correlation parameter by weighted sum averaging, so as to realize the comprehensive quantification of different cancer types in two dimensions of patient overlap and feature similarity, which helps to accurately model the potential correlation relationship between cancer types, provides a more discriminative correlation input for the risk prediction model in the multi-cancer classification scenario, and improves the model's ability to identify boundary samples and the overall prediction accuracy.
[0102] As an optional embodiment, in the above steps, optimizing and training a prediction model for predicting the cancer risk corresponding to any cancer type according to the type correlation parameter and the data feature set corresponding to each cancer type includes:
[0103] The data feature set corresponding to each cancer type is used as the training data labeled with the corresponding cancer type and input into the basic model corresponding to each cancer type for training. During the training process, all basic models corresponding to different cancer types are jointly optimized and trained based on a common loss function until convergence, thereby obtaining the prediction model corresponding to each cancer type. Optionally, the common loss function is the weighted sum average of the loss function values of each basic model, where the weight of the loss function value corresponding to each basic model is proportional to the average value of all type correlation parameters corresponding to the cancer type corresponding to the basic model.
[0104] The prediction models corresponding to all cancer types are determined as the common prediction model, which is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any cancer type.
[0105] It can be seen that through the above optional embodiments, a basic model is constructed based on the data feature set corresponding to each cancer type, and a common loss function is introduced during the training process to uniformly optimize all basic models, enabling the training process to simultaneously consider the correlations between different cancer types. The weight of each model in the overall loss is dynamically adjusted according to its correlation degree, thereby enhancing the expression ability of models for weakly related cancer types and effectively enhancing the generalization performance and collaborative prediction ability of the overall model in the multi-cancer classification task. Finally, a unified common prediction model for multi-cancer type risk assessment is formed, achieving disease prediction with higher accuracy and stronger adaptability.
[0106] Embodiment 2
[0107] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a training system for a cancer risk prediction model based on joint optimization according to an embodiment of the present invention. Among them, Figure 2 the described training system for a cancer risk prediction model based on joint optimization can be applied to a data processing system / data processing device / data processing server (wherein, the server includes a local processing server or a cloud processing server). As Figure 2 shown, the training system for a cancer risk prediction model based on joint optimization may include:
[0108] An acquisition module 201, configured to acquire cfDNA sequencing data and corresponding methylation characteristics of multiple patients with multiple cancer types.
[0109] A first determination module 202, configured to determine the data feature set corresponding to each cancer type based on a classification algorithm according to the cfDNA sequencing data and the corresponding methylation characteristics.
[0110] A second determination module 203, configured to determine the type correlation parameter between any two cancer types.
[0111] A training module 204, configured to optimize and train a prediction model for predicting the cancer risk corresponding to any cancer type according to the type correlation parameter and the data feature set corresponding to each cancer type.
[0112] It can be seen that the above-described inventive embodiments obtain the sequencing data and corresponding methylation characteristics of patients with multiple cancer types, extract the characteristic patterns of various cancers through a classification algorithm, and optimize and train in combination with the correlation parameters between different cancer types to obtain a prediction model for predicting the cancer risk corresponding to any of the cancer types. Thereby, the ability of the model to distinguish different cancer types and the generalization ability can be improved, the prediction accuracy and adaptability of the cancer risk corresponding to any cancer type can be enhanced, and a more scientific and reliable model support can be provided for individualized cancer screening and early intervention.
[0113] As an optional embodiment, the methylation characteristics include the methylation site characteristics of at least one region of the cfDNA sequencing data; the methylation site characteristics include at least one of a methylation rate, a continuous methylation site characteristic, a methylation site change characteristic, and a methylation site map characteristic.
[0114] It can be seen that through the above optional embodiment, the content of the methylation characteristics is defined to comprehensively characterize the methylation-related characteristics of the patient, so as to facilitate subsequent type correlation calculation and model training, and assist in improving the ability of the model to distinguish different cancer types and the generalization ability, enhancing the prediction accuracy and adaptability of the cancer risk corresponding to any cancer type, and providing a more scientific and reliable model support for individualized cancer screening and early intervention.
[0115] As an optional embodiment, the cancer types include at least one of liver cancer, lung cancer, gastric cancer, colorectal cancer, breast cancer, and nasopharyngeal cancer.
[0116] It can be seen that through the above optional embodiment, the types of cancer types are defined to facilitate subsequent type correlation calculation and model training, and assist in improving the ability of the model to distinguish different cancer types and the generalization ability, enhancing the prediction accuracy and adaptability of the cancer risk corresponding to any cancer type, and providing a more scientific and reliable model support for individualized cancer screening and early intervention.
[0117] As an optional embodiment, the specific manner in which the first determination module determines the data feature set corresponding to each cancer type based on a classification algorithm according to the cfDNA sequencing data and the corresponding methylation characteristics includes:
[0118] For each patient, when the cancer type of the patient is a single cancer type, the cfDNA sequencing data and the corresponding methylation characteristics of the patient are classified into the data feature set corresponding to the corresponding cancer type;
[0119] When the cancer type of the patient is a multi-cancer type, based on the patient's medical history, cfDNA sequencing data, and corresponding methylation characteristics, determine the closest cancer type corresponding to the patient, and classify the patient's cfDNA sequencing data and corresponding methylation characteristics into the data feature set of the closest cancer type.
[0120] It can be seen that through the above optional embodiments, the cancer type of each patient is judged. When it is a single cancer type, directly classify its cfDNA sequencing data and methylation characteristics into the data feature set of the corresponding type. When it is a multi-cancer type, then combine its medical history and methylation characteristics to determine the closest cancer type, and classify the relevant data into the data feature set of this type, thereby improving the classification accuracy and biological rationality when constructing the data feature set, helping to enhance the adaptability and generalization ability of the subsequent prediction model to complex clinical situations, and improving the scientificity and stability of multi-cancer type risk prediction.
[0121] As an optional embodiment, the specific way for the first determination module to determine the closest cancer type corresponding to the patient based on the patient's medical history, cfDNA sequencing data, and corresponding methylation characteristics includes:
[0122] Determine the reference feature data corresponding to each cancer type included in the cancer type of the patient;
[0123] Calculate the similarity between the patient's cfDNA sequencing data and corresponding methylation characteristics and the reference feature data corresponding to each cancer type to obtain the similarity parameter corresponding to each cancer type;
[0124] Calculate the confirmed diagnosis duration corresponding to each cancer type in the patient's medical history, and calculate the duration weight proportional to the confirmed diagnosis duration corresponding to each cancer type;
[0125] Calculate the product of the similarity parameter and the duration weight corresponding to each cancer type;
[0126] Determine the cancer type with the highest product as the closest cancer type corresponding to the patient.
[0127] It can be seen that through the above optional embodiments, for multi-cancer type patients, by obtaining the reference feature data of their corresponding cancer types, calculating the similarity between their cfDNA sequencing data and methylation characteristics and each cancer type respectively, and generating a duration weight in combination with the confirmed diagnosis duration information in the medical history, further determining the closest cancer type based on the product of the similarity parameter and the duration weight, so as to achieve the accurate classification of multi-cancer type patients, effectively improve the accuracy of subsequent data feature set construction, and enhance the recognition ability and classification effect of the cancer risk prediction model for multi-source heterogeneous samples.
[0128] As an alternative embodiment, the specific manner in which the second determination module determines the type correlation parameter between any two cancer types includes:
[0129] Obtain the historical patient data corresponding to each cancer type;
[0130] For any two cancer types, determine the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types.
[0131] It can be seen that through the above alternative embodiment, based on the historical patient data of each cancer type and combined with its corresponding data feature set, the type correlation parameter between any two cancer types is calculated, so as to accurately reflect the similarity between different cancer types at the methylation feature level, effectively improve the discrimination ability of cancer classification, provide an accurate basis for introducing type correlation information into the subsequent prediction model, thereby optimizing the model structure and improving the accuracy and generalization ability of cancer risk prediction.
[0132] As an alternative embodiment, the specific manner in which the second determination module determines the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types includes:
[0133] For any two cancer types, calculate the proportion of the number of identical patients between the historical patient data of the two cancer types;
[0134] Calculate the data similarity between the data feature sets corresponding to the two cancer types;
[0135] Calculate the weighted sum average of the proportion of the number of identical patients and the data similarity to obtain the type correlation parameter between the two cancer types.
[0136] It can be seen that through the above alternative embodiment, based on the historical patient data between any two cancer types, the proportion of the number of identical patients is calculated, and combined with the similarity between the corresponding data feature sets, the type correlation parameter is determined by the weighted sum average method, so as to realize the comprehensive quantification of different cancer types in two dimensions of patient overlap and feature similarity, which helps to accurately model the potential correlation relationship between cancer types, provides a more discriminative correlation input for the risk prediction model in the multi-cancer classification scenario, and improves the model's ability to identify boundary samples and the overall prediction accuracy.
[0137] As an alternative embodiment, the specific manner in which the training module optimizes and trains a prediction model for predicting the cancer risk corresponding to any cancer type according to the type correlation parameter and the data feature set corresponding to each cancer type includes:
[0138] The data feature set corresponding to each cancer type is used as the training data labeled with the corresponding cancer type and input into the basic model corresponding to each cancer type for training. During the training process, all basic models corresponding to different cancer types are jointly optimized and trained based on a common loss function until convergence, and a prediction model corresponding to each cancer type is obtained. Optionally, the common loss function is the weighted sum average of the loss function values of each basic model, where the weight of the loss function value corresponding to each basic model is proportional to the average value of all type correlation parameters corresponding to the cancer type corresponding to the basic model.
[0139] The prediction models corresponding to all cancer types are determined as the common prediction model, and the common prediction model is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any cancer type.
[0140] It can be seen that through the above optional embodiments, a basic model is constructed based on the data feature set corresponding to each cancer type, and a common loss function is introduced during the training process to uniformly optimize all basic models, enabling the training process to consider the correlation between different cancer types simultaneously. The weight of each model in the overall loss is dynamically adjusted according to its correlation degree, thereby improving the expression ability of the models for weakly related cancer types, effectively enhancing the generalization performance and collaborative prediction ability of the overall model in the multi-cancer classification task, and finally forming a unified common prediction model that can be used for multi-cancer type risk assessment, achieving disease prediction with higher accuracy and stronger adaptability.
[0141] Embodiment III
[0142] Please refer to Figure 3 , Figure 3 which is another training system for a cancer risk prediction model based on joint optimization disclosed in the embodiments of the present invention. Figure 3 The described training system for a cancer risk prediction model based on joint optimization is applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). As Figure 3 shown, the training system for a cancer risk prediction model based on joint optimization may include:
[0143] A memory 301 storing executable program code;
[0144] A processor 302 coupled to the memory 301;
[0145] Wherein, the processor 302 calls the executable program code stored in the memory 301 to execute the steps of the training method for a cancer risk prediction model based on joint optimization described in Embodiment I.
[0146] Embodiment IV
[0147] An embodiment of the present invention discloses a computer-readable storage medium that stores a computer program for electronic data exchange. Among them, the computer program enables a computer to execute the steps of the training method of the cancer risk prediction model based on co-optimization described in the first embodiment.
[0148] Embodiment Five
[0149] An embodiment of the present invention discloses a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the training method of the cancer risk prediction model based on co-optimization described in the first embodiment.
[0150] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or continuous order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0151] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0152] For convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0153] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0155] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 block or multiple blocks.
[0157] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0158] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0159] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0160] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0161] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0162] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0163] Finally, it should be noted that what is disclosed by a training method and system of a cancer risk prediction model based on co-optimization according to the embodiments of the present invention is only the preferred embodiments of the present invention, which are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a cancer risk prediction model based on co-optimization, characterized in that The method includes: Obtaining cfDNA sequencing data and corresponding methylation characteristics of multiple patients with multiple cancer types; For each of the patients, when the cancer type of the patient is a single cancer type, classifying the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set corresponding to the cancer type; When the cancer type of the patient is a multi-cancer type, determining the reference feature data corresponding to each cancer type included in the cancer type of the patient; Calculating the similarity between the cfDNA sequencing data and the corresponding methylation characteristics of the patient and the reference feature data corresponding to each cancer type to obtain a similarity parameter corresponding to each cancer type; Calculating the duration of disease diagnosis corresponding to each cancer type in the patient's disease history, and calculating a duration weight proportional to the duration of disease diagnosis corresponding to each cancer type; Calculating the product of the similarity parameter corresponding to each cancer type and the duration weight; Determining the cancer type with the highest product as the closest cancer type corresponding to the patient, and classifying the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set of the closest cancer type; Determining the type correlation parameter between any two of the cancer types; Taking the data feature set corresponding to each cancer type as training data labeled with the corresponding cancer type, and inputting it into the basic model corresponding to each cancer type for training. During the training process, jointly optimizing and training the basic models corresponding to all cancer types based on a common loss function until convergence to obtain a prediction model corresponding to each cancer type; the common loss function is the weighted sum average of the loss function values of each basic model, where the weight of the loss function value corresponding to each basic model is proportional to the average value of all the type correlation parameters corresponding to the cancer type corresponding to the basic model; Determining all the prediction models corresponding to the cancer types as a common prediction model, and the common prediction model is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any one of the cancer types.
2. The training method of the cancer risk prediction model based on co-optimization according to claim 1, wherein The methylation characteristics include methylation site characteristics of at least one region of the cfDNA sequencing data; the methylation site characteristics include at least one of methylation rate, methylation site continuous characteristics, methylation site change characteristics, and methylation site map characteristics.
3. The training method of the cancer risk prediction model based on co-optimization according to claim 1, characterized in that, The cancer types include at least one of liver cancer, lung cancer, gastric cancer, colorectal cancer, breast cancer, and nasopharyngeal cancer.
4. The training method of the cancer risk prediction model based on co-optimization according to claim 1, wherein, The determining the type correlation parameter between any two of the cancer types includes: Obtaining the historical patient data corresponding to each cancer type; For any two of the cancer types, determining the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types.
5. The training method of the cancer risk prediction model based on joint optimization according to claim 4, characterized in that, The determining the type correlation parameter between the two cancer types according to the historical patient data and the data feature set of the two cancer types includes: For any two of the cancer types, calculate the proportion of the number of identical patients between the historical patient data of the two cancer types; Calculate the data similarity between the data feature sets corresponding to the two cancer types; Calculate the weighted sum average of the proportion of the number of identical patients and the data similarity to obtain the type correlation parameter between the two cancer types.
6. A training system for a cancer risk prediction model based on co-optimization, characterized in that, The system includes: An acquisition module for acquiring cfDNA sequencing data and corresponding methylation characteristics of multiple patients with multiple cancer types; A first determination module for determining, based on a classification algorithm, the data feature set corresponding to each cancer type according to the cfDNA sequencing data and the corresponding methylation characteristics, specifically including: For each patient, when the cancer type of the patient is a single cancer type, classify the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set of the corresponding cancer type; When the cancer type of the patient is a multi-cancer type, determine the reference feature data corresponding to each cancer type included in the cancer type of the patient; Calculate the similarity between the cfDNA sequencing data and the corresponding methylation characteristics of the patient and the reference feature data corresponding to each cancer type to obtain the similarity parameter corresponding to each cancer type; Calculate the duration of disease diagnosis corresponding to each cancer type in the patient's disease history, and calculate the duration weight proportional to the duration of disease diagnosis corresponding to each cancer type; Calculate the product of the similarity parameter corresponding to each cancer type and the duration weight; Determine the cancer type with the highest product as the closest cancer type corresponding to the patient, and classify the cfDNA sequencing data and the corresponding methylation characteristics of the patient into the data feature set of the closest cancer type; A second determination module for determining the type correlation parameter between any two of the cancer types; A training module for optimizing and training, according to the type correlation parameter and the data feature set corresponding to each cancer type, to obtain a prediction model for predicting the cancer risk corresponding to any cancer type, specifically including: Use the data feature set corresponding to each cancer type as training data labeled with the corresponding cancer type, and input it into the basic model corresponding to each cancer type for training. During the training process, jointly optimize and train the basic models corresponding to all cancer types based on a common loss function until convergence to obtain the prediction model corresponding to each cancer type; the common loss function is the weighted sum average of the loss function values of each basic model, where the weight of the loss function value corresponding to each basic model is proportional to the average value of all the type correlation parameters corresponding to the cancer type corresponding to the basic model; Determine all the prediction models corresponding to the cancer types as a common prediction model, and the common prediction model is used to receive the data to be predicted of a patient to predict the cancer risk corresponding to any cancer type.
7. A training system for a cancer risk prediction model based on co-optimization, characterized in that, The system includes: A memory storing executable program code; A processor coupled to the memory; The processor invokes the executable program code stored in the memory and executes the training method based on the jointly optimized cancer risk prediction model according to any one of claims 1-5.
Citation Information
Patent Citations
Cancer survival prediction model construction method based on graph contrast learning
CN115985442A
Method for predicting distant metastasis state of colorectal cancer
CN118538416A