Method and system for constructing cancer early screening model based on difference statistics
By combining differential statistics and tree models, the target category clusters are selected through clustering, confusing features are eliminated, and feature combinations are optimized. This solves the problem of difficult feature combination evaluation in traditional methods and achieves high accuracy and stability of the early cancer screening model.
Patent Information
- Application Number
- CN202510972908.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-07-15
Smart Images

Figure CN120853693A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cancer model training technology, specifically to a method and system for constructing a cancer early screening model based on differential statistics. Background Technology
[0002] Within living organisms, during processes such as apoptosis, DNA fragments within cells are released into the blood plasma, becoming cell-free DNA (cfDNA). In the early stages of cancer development, before patients exhibit obvious clinical symptoms, the state of intracellular DNA has already changed, and this DNA is released into the blood plasma. This results in plasma cfDNA containing cancer-related information. By extracting and processing this information, non-invasive cancer diagnosis can be performed, enabling early detection and treatment.
[0003] The current mainstream approach to cancer cfDNA research is to infer cancer occurrence by utilizing mutations in single or a few cancer-related genes on cfDNA. In early cancer screening, the effective combination of cfDNA fragment omics features (including fragment size, fragment terminal base sequence, fragment breakpoint base sequence, and nucleosome footprint, etc.) is crucial for improving screening accuracy.
[0004] Cancer data typically exhibits high dimensionality, meaning the number of features far exceeds the number of samples. This significantly increases model complexity, leading to computational costs, the curse of dimensionality, overfitting, and negatively impacting generalization ability. Feature selection, crucial for cancer data analysis, aims to identify the most informative variables from high-dimensional data, eliminating redundancy and reducing noise. Effective feature selection methods can improve classifier performance and survival analysis model prediction accuracy, and also uncover disease-related biomarkers, providing more reliable data for medical decision-making.
[0005] However, there is a lack of intuitive and effective means to evaluate high-dimensional feature combinations. In existing technologies, traditional feature selection methods struggle to visually demonstrate the impact of feature combinations on sample classification, and cannot quickly determine the effectiveness of feature combinations from a visualization perspective. While the t-SNE algorithm can be used for high-dimensional data visualization, its application in cfDNA fragmentomics feature combination evaluation lacks a systematic process and optimization strategies specific to this field, resulting in an inability to accurately assess the actual value of feature combinations in early cancer screening. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for constructing a cancer early screening model based on differential statistics, thereby solving the following technical problems: Traditional feature selection methods are difficult to intuitively demonstrate the impact of feature combinations on sample classification, and cannot quickly determine the effectiveness of feature combinations from a visual perspective.
[0007] The objective of this invention can be achieved through the following technical solutions: A method for constructing a cancer early screening model based on differential statistics includes the following steps: S1, Obtain normal sample group and cancer sample group, wherein the normal sample group and cancer sample group each contain multiple sample data, and extract the corresponding original feature set of circulating cell-free DNA for each sample data; S2, integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples, perform clustering on the joint dataset to obtain several category clusters; count the proportion of cancer samples in any category cluster, sort all category clusters according to the proportion, select the top N category clusters as target category clusters, extract the dataset of all normal samples from any target category cluster and label it as the target sample group; N is a preset quantity threshold; S3, remove the target sample group from the normal sample group to obtain the comparison sample group, input the target sample group and the comparison sample group into the preset tree model to obtain the importance score of any original feature, and label the original features with an importance score greater than or equal to a preset threshold as undetermined features to obtain the first undetermined feature set; input the cancer sample group and the comparison sample group into the preset tree model and repeat the above process to obtain the second undetermined feature set; S4, obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and construct an initial classification model based on the target feature set; S5. Based on the initial classification model, perform incremental training on each common feature in the common feature set to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. Determine the performance score of all incremental models based on the best-in-best solution distance method, and select the incremental model with the largest performance score as the final model.
[0008] As a further aspect of the present invention: in S1, the original feature set includes fragment size distribution parameters, terminal base sequence fundamental frequency features, breakpoint base sequence fundamental frequency features, and nucleosome footprint features.
[0009] As a further aspect of the present invention: S2 further includes removing a target category cluster if the total number of samples in any target cluster is less than a preset total number of samples.
[0010] As a further aspect of the present invention: in step S3, the specific process for obtaining the important scores is as follows: The target sample group and the comparison sample group are labeled with different group labels, and together with the original feature data, they form a labeled dataset. The labeled dataset is input into a preset tree model. According to the splitting rules of the tree model, the amount of impurity reduction caused by the splitting of each original feature at the tree node is calculated. The amount of impurity reduction of the same original feature in all trees is accumulated and normalized. The data are then sorted in descending order according to the numerical value to obtain the importance score corresponding to each original feature.
[0011] As a further aspect of the present invention, the specific construction process of the initial classification model is as follows: A feature filtering operation is performed on the joint dataset, retaining only the original features in the target feature set and removing all other original features. The filtered joint dataset is then divided into a training set and a test set according to a preset ratio. A preset general model is trained based on the training set to obtain an initial classification model.
[0012] As a further aspect of the present invention: S5 further includes obtaining the performance score corresponding to the initial classification model; if the performance score of any incremental model is greater than the performance score corresponding to the initial classification model, then the common feature corresponding to the incremental model is retained; if the performance score of any incremental model is less than or equal to the performance score corresponding to the initial classification model, then the common feature corresponding to the incremental model is removed.
[0013] As a further aspect of the present invention: the common features retained are combined using an exhaustive method to generate several experimental groups. Incremental training is performed on each experimental group based on the initial classification model to obtain the incremental model and performance score corresponding to each experimental group. The incremental model with the highest performance score is selected as the final model.
[0014] As a further aspect of the present invention: if the performance scores of all incremental models are less than or equal to the performance score corresponding to the initial classification model, then the initial classification model is directly used as the final model.
[0015] The present invention also includes a system for constructing a cancer early screening model based on differential statistics, for implementing the above-described method for constructing a cancer early screening model based on differential statistics, comprising: The data acquisition module is used to acquire normal sample groups and cancer sample groups, wherein each normal sample group and cancer sample group contains multiple sample data, and the corresponding original feature set of circulating cell-free DNA is extracted for each sample data. The data analysis module is used to integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples. Clustering is performed on the joint dataset to obtain several class clusters. The proportion of cancer samples in any class cluster is counted, and all class clusters are sorted according to the proportion. The top N class clusters are selected as target class clusters. The dataset of all normal samples is extracted from any target class cluster and labeled as the target sample group. N is a preset quantity threshold. The data optimization module is used to remove the target sample group from the normal sample group to obtain the comparison sample group. The target sample group and the comparison sample group are input into a preset tree model to obtain the importance score of any original feature. The original features with an importance score greater than or equal to a preset threshold are labeled as undetermined features to obtain the first undetermined feature set. The cancer sample group and the comparison sample group are input into the preset tree model, and the above process is repeated to obtain the second undetermined feature set. The model building module is used to obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and build an initial classification model based on the target feature set; The model optimization module is used to perform incremental training on each common feature in the common feature set based on the initial classification model, to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. The performance score of all incremental models is determined based on the best-in-best solution distance method, and the incremental model with the largest performance score is selected as the final model.
[0016] The beneficial effects of this invention are: 1) By clustering the joint dataset of normal and cancer samples, target clusters are selected according to the proportion of cancer samples. The extracted "abnormal" normal samples (i.e., normal samples with features similar to cancer samples) can pinpoint the first original feature set causing feature confusion. This process effectively identifies feature confusion points in potential false negative samples, avoiding misjudgments of cancer due to abnormal features in normal samples. It improves the model's ability to distinguish between true and false positives from the sample selection stage, solving the problem of insufficient early screening specificity caused by feature confusion in traditional methods, and laying a precise data foundation for subsequent feature removal and model optimization.
[0017] 2) From the second original feature set obtained based on clustering and tree models, the first original feature set, which contains features that cause confusion, is selectively removed. This directly eliminates redundant features that lead to ambiguity in sample classification. This feature removal strategy based on differential statistics can retain key features unique to cancer samples (such as fragment size distribution and terminal base sequence) while avoiding interference from feature confusion in model classification. This allows the model to focus more on cancer-specific features, effectively reduces the impact of noise in high-dimensional data, improves the diagnostic efficacy of the feature set for cancer samples, and provides a cleaner feature input for early screening models.
[0018] 3) After constructing the initial model by removing confusing features, incremental training is performed on the common feature set. This is then evaluated using multiple metrics such as accuracy, AUC, and the distance between best and worst solutions. This process can uncover potentially useful features masked by confusing features in the second original feature set. By dynamically evaluating the impact of feature addition on model performance, this avoids directly removing potentially useful features and improves the model's generalization ability by gradually optimizing feature combinations. It overcomes the "either / or" limitation of traditional feature selection, enabling refined feature mining for early cancer screening and ultimately forming an optimal feature combination model that balances specificity and sensitivity.
[0019] 4) During incremental training and feature combination optimization, exhaustive methods are used to combine retained features and capture nonlinear interactions between features (such as the synergistic effect of fragment size distribution and nucleosome footprint features) based on the splitting rules of the tree model. This avoids the loss of interaction information caused by the independent evaluation of feature importance in traditional feature screening, thereby effectively mining the combined effects of cancer-related features. For example, some base motif features only show diagnostic value under specific fragment size conditions, so that the model can more comprehensively characterize the complex relationship between cfDNA features and cancer status, further improving the accuracy and robustness of the early screening model. Attached Figure Description
[0020] The invention will now be further described with reference to the accompanying drawings.
[0021] Figure 1 This is a schematic diagram of the construction method of a cancer early screening model based on differential statistics according to the present invention. Detailed Implementation
[0022] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0023] Please see Figure 1As shown, this invention provides a method for constructing a cancer early screening model based on differential statistics, comprising the following steps: S1, Obtain normal sample group and cancer sample group, wherein the normal sample group and cancer sample group each contain multiple sample data, and extract the corresponding original feature set of circulating cell-free DNA for each sample data; S2, integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples, perform clustering on the joint dataset to obtain several category clusters; count the proportion of cancer samples in any category cluster, sort all category clusters according to the proportion, select the top N category clusters as target category clusters, extract the dataset of all normal samples from any target category cluster and label it as the target sample group; N is a preset quantity threshold; S3, remove the target sample group from the normal sample group to obtain the comparison sample group, input the target sample group and the comparison sample group into the preset tree model to obtain the importance score of any original feature, and label the original features with an importance score greater than or equal to a preset threshold as undetermined features to obtain the first undetermined feature set; input the cancer sample group and the comparison sample group into the preset tree model and repeat the above process to obtain the second undetermined feature set; S4, obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and construct an initial classification model based on the target feature set; S5. Based on the initial classification model, perform incremental training on each common feature in the common feature set to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. Determine the performance score of all incremental models based on the best-in-best solution distance method, and select the incremental model with the largest performance score as the final model.
[0024] 1) By clustering the joint dataset of normal and cancer samples, target clusters are selected according to the proportion of cancer samples. The extracted "abnormal" normal samples (i.e., normal samples with features similar to cancer samples) can pinpoint the first original feature set causing feature confusion. This process effectively identifies feature confusion points in potential false negative samples, avoiding misjudgments of cancer due to abnormal features in normal samples. It improves the model's ability to distinguish between true and false positives from the sample selection stage, solving the problem of insufficient early screening specificity caused by feature confusion in traditional methods, and laying a precise data foundation for subsequent feature removal and model optimization.
[0025] 2) From the second original feature set obtained based on clustering and tree models, the first original feature set, which contains features that cause confusion, is selectively removed. This directly eliminates redundant features that lead to ambiguity in sample classification. This feature removal strategy based on differential statistics can retain key features unique to cancer samples (such as fragment size distribution and terminal base sequence) while avoiding interference from feature confusion in model classification. This allows the model to focus more on cancer-specific features, effectively reduces the impact of noise in high-dimensional data, improves the diagnostic efficacy of the feature set for cancer samples, and provides a cleaner feature input for early screening models.
[0026] 3) After constructing the initial model by removing confusing features, incremental training is performed on the common feature set. This is then evaluated using multiple metrics such as accuracy, AUC, and the distance between best and worst solutions. This process can uncover potentially useful features masked by confusing features in the second original feature set. By dynamically evaluating the impact of feature addition on model performance, this avoids directly removing potentially useful features and improves the model's generalization ability by gradually optimizing feature combinations. It overcomes the "either / or" limitation of traditional feature selection, enabling refined feature mining for early cancer screening and ultimately forming an optimal feature combination model that balances specificity and sensitivity.
[0027] 4) During incremental training and feature combination optimization, exhaustive methods are used to combine retained features and capture nonlinear interactions between features (such as the synergistic effect of fragment size distribution and nucleosome footprint features) based on the splitting rules of the tree model. This avoids the loss of interaction information caused by the independent evaluation of feature importance in traditional feature screening, thereby effectively mining the combined effects of cancer-related features. For example, some base motif features only show diagnostic value under specific fragment size conditions, so that the model can more comprehensively characterize the complex relationship between cfDNA features and cancer status, further improving the accuracy and robustness of the early screening model.
[0028] In a preferred embodiment of the present invention, in S1, the original feature set includes fragment size distribution parameters, terminal base sequence fundamental frequency features, breakpoint base sequence fundamental frequency features, and nucleosome footprint features.
[0029] In another preferred embodiment of the present invention, step S2 further includes removing a target cluster if the total number of samples in any target cluster is less than a preset total number of samples.
[0030] After performing clustering on the joint dataset, several class clusters are obtained. For each class cluster, the total number of samples it contains is counted. A preset threshold for the total number of samples is set to ensure that there are enough samples within the class cluster to reflect the characteristic distribution pattern. For example, according to common statistical sampling principles, if the sample size is too small, the characteristic distribution may be random and cannot represent the overall situation. When the total number of samples in a target class cluster is less than the preset threshold, it indicates that the number of samples in that cluster is insufficient and cannot reliably reflect the characteristic pattern of that class cluster, so it is removed.
[0031] It is crucial to ensure that the retained target category clusters have a sufficient sample size, thereby making the target sample groups extracted from the target category clusters more representative. Only with a sufficient sample size can the feature distribution of the category clusters be guaranteed to be stable and not formed by chance. This helps in achieving the ultimate goal of the scheme by avoiding bias in the extracted target sample groups due to insufficient sample size in the category clusters, which would affect the accuracy of subsequent feature importance analysis. This ensures that the cancer early screening model built based on these samples and features can more reliably identify cancer samples from normal samples, improving the model's stability and accuracy.
[0032] In another preferred embodiment of the present invention, the specific process of obtaining the important score in S3 is as follows: The target sample group and the comparison sample group are labeled with different group labels, and together with the original feature data, they form a labeled dataset. The labeled dataset is input into a preset tree model. According to the splitting rules of the tree model, the amount of impurity reduction caused by the splitting of each original feature at the tree node is calculated. The amount of impurity reduction of the same original feature in all trees is accumulated and normalized. The data are then sorted in descending order according to the numerical value to obtain the importance score corresponding to each original feature.
[0033] First, the target sample group and the comparison sample group are assigned different group labels, such as "Group A" for the target sample group and "Group B" for the comparison sample group. These labels are used to distinguish the attributes of the two types of samples. Then, these group labels are combined with the original feature data corresponding to each sample (such as fragment size distribution parameters, terminal base sequence fundamental frequency features, etc.) to form a labeled dataset. This allows the tree model to determine which group each sample's features belong to. Next, the labeled dataset is input into the pre-defined tree model. During the tree model's construction, node splitting is performed based on the feature's ability to classify samples. Each split aims to reduce the impurity of samples within the resulting child nodes. For each original feature, when it is used for tree node splitting, the amount of impurity reduction caused by this split is calculated. For example, if a feature splits a node, the categories of samples within the child nodes become more concentrated, and the impurity is reduced to a certain extent. Since a tree model may consist of multiple trees, it is necessary to sum up the amount of impurity reduction generated by the same original feature in each split across all trees, then perform normalization, and finally sort them in descending order of importance scores to obtain the importance score of each original feature.
[0034] By utilizing the splitting mechanism of a tree model, the importance of each original feature in distinguishing the target sample group from the control sample group is objectively quantified. The advantage lies in leveraging the inherent suitability of tree models for handling high-dimensional features and their ability to automatically assess feature importance, avoiding subjective judgments of feature value and making the feature selection process more scientific. Ultimately, after accurately obtaining feature importance scores, the system can identify features more critical to distinguishing the "abnormal" group (target sample group) from ordinary normal samples (control sample group). These features may reflect aspects of normal samples that are confused with cancer sample features, laying the foundation for subsequent removal of confusing features and retention of truly distinguishable features. This allows the constructed early cancer screening model to more accurately identify sample categories, reduce misjudgments caused by feature confusion, and improve the accuracy and reliability of early screening.
[0035] In another preferred embodiment of the present invention, the specific construction process of the initial classification model is as follows: A feature filtering operation is performed on the joint dataset, retaining only the original features in the target feature set and removing all other original features. The filtered joint dataset is then divided into a training set and a test set according to a preset ratio. A preset general model is trained based on the training set to obtain an initial classification model.
[0036] When performing feature filtering on a joint dataset, the original features in the target feature set are first defined, for example, assuming the target feature set contains fragment size distribution parameters, specific terminal base sequence fundamental frequency features, etc. Then, each sample in the joint dataset is iterated over. For each sample's feature data, only the features from the target feature set are retained, and all other original features not in the target feature set are removed. For example, if a sample originally has 100 original features, of which 20 belong to the target feature set, only these 20 features are retained, and the remaining 80 are removed. Next, the filtered joint dataset is divided into training and test sets according to a preset ratio (such as the common 7:3 or 8:2). Random sampling is usually used during this division to ensure that the sample distribution of the training and test sets is similar to that of the original dataset. Then, a preset general model (such as a logistic regression model or a support vector machine model) is trained on the training set. During training, the model learns the mapping relationship between the target features and the sample category (normal or cancer). By continuously adjusting the model parameters, the model can better fit the data on the training set, ultimately obtaining the initial classification model. For example, a logistic regression model can be trained using 70% of the filtered data, allowing the model to learn to determine the sample category based on the retained features, resulting in a preliminary classification model.
[0037] Feature filtering reduces interference from redundant features, allowing the model to focus on learning the more critical features in the target feature set that distinguish cancerous and normal samples, thus improving training efficiency and generalization ability. Dividing the model into training and test sets allows for performance evaluation using the unused test set after training, ensuring the model has not simply memorized the training data but truly learned the characteristic patterns of the samples. Training a general model based on the training set yields an initial classification model, providing a foundation for subsequent incremental training and model optimization. This contributes to the ultimate goal of the solution by reducing feature dimensionality, lowering computational costs, avoiding the curse of dimensionality and overfitting problems associated with high-dimensional features, and enabling the initial classification model to more accurately classify samples by focusing on key features. Evaluation on the test set ensures the model's applicability to new data, laying a solid foundation for further performance optimization and building a more accurate early cancer screening model.
[0038] In another preferred embodiment of the present invention, step S5 further includes obtaining the performance score corresponding to the initial classification model; if the performance score of any incremental model is greater than the performance score corresponding to the initial classification model, then the common feature corresponding to the incremental model is retained; if the performance score of any incremental model is less than or equal to the performance score corresponding to the initial classification model, then the common feature corresponding to the incremental model is removed.
[0039] After obtaining the initial classification model, its performance score needs to be calculated. Specifically, the initial model is evaluated using a pre-defined test set. Performance metrics such as accuracy, AUC, and loss on the test set are calculated, and then the Best-to-Best Solution Distance (TOPSIS) method is used to synthesize these metrics into a performance score. For example, if the initial model has an accuracy of 80%, an AUC of 0.85, and a loss of 0.3 on the test set, a comprehensive score is obtained after TOPSIS processing. Next, incremental training is performed on each feature in the common feature set. Each time, a common feature is added to the initial model, and the model is retrained to obtain an incremental model. The performance metrics of each incremental model are then evaluated using the test set and converted into a corresponding performance score. For example, if the common feature set contains features A and B, feature A is added first to train incremental model 1, and a score is obtained after evaluation; then feature B is added to train incremental model 2, and a different score is obtained after evaluation. Then, the performance score of each incremental model is compared with the performance score of the initial model. If the score of an incremental model is higher than that of the initial model, it means that adding this common feature improves the model's performance, so the feature is retained. If the score of an incremental model is less than or equal to that of the initial model, it means that adding this feature does not improve the model's performance or even decreases it, so the feature is removed. For example, if the score of incremental model 1 is higher than that of the initial model, feature A is retained; if the score of incremental model 2 is equal to that of the initial model, feature B is removed.
[0040] By comparing the performance of the incremental model with the initial model, the actual value of each feature in the common feature set to the model can be determined, thus selecting features that truly improve model performance. The advantage is that it avoids blindly retaining all common features; by using quantitative performance evaluation to decide feature retention, it ensures that the ultimately retained features are those that positively impact the model, avoiding the introduction of redundant or ineffective features. This contributes to the ultimate goal of the solution: features selected in this way can optimize the model's feature combination, enabling the model to more accurately identify cancer samples in subsequent training and applications, improving the model's classification performance and generalization ability. This leads to the construction of a more accurate and reliable early cancer screening model, enhancing the accuracy and effectiveness of early cancer screening, and providing stronger support for early cancer diagnosis and treatment.
[0041] In another preferred embodiment of the present invention, an exhaustive method is used to combine the retained common features to generate several experimental groups. Incremental training is performed on each experimental group based on the initial classification model to obtain the incremental model and performance score corresponding to each experimental group. The incremental model with the largest performance score is selected as the final model.
[0042] After initially filtering and retaining some common features by comparing the performance scores of the incremental model and the initial model, an exhaustive method is used to combine these retained common features. The principle of the exhaustive method is to perform full permutations and combinations of all retained features, generating all possible feature combinations. For example, if features A, B, and C are retained, the exhaustive method will generate different experimental groups such as {A}, {B}, {C}, {A,B}, {A,C}, {B,C}, and {A,B,C}. After generating several experimental groups, incremental training is performed on each experimental group based on the initial classification model. This involves adding a feature combination from one experimental group to the initial model each time and retraining the model using the training set, resulting in an incremental model for each experimental group. Then, each incremental model is evaluated using a test set, calculating its accuracy, AUC value, and loss value, and these metrics are then combined using the best-case distance method to convert them into a performance score for each incremental model. Finally, the performance scores of all incremental models are compared, and the incremental model with the highest performance score is selected as the final model. For example, if the incremental model corresponding to the experimental group {A,B} has the highest performance score, then that model will be selected as the final model for early cancer screening.
[0043] By exhaustively exploring all possible combinations of features, the most significant feature combination for improving model performance can be found. Because the improvement of a single feature may be limited, while different features may interact, combining different features may produce a synergistic effect, further improving the model's classification performance. The advantage of exhaustive exploration is that it ensures no possible effective feature combination is overlooked. By systematically evaluating each combination, the optimal feature combination is determined quantitatively, avoiding the limitations of selecting feature combinations based on experience or subjective judgment. This contributes to the ultimate goal of the approach: by finding the optimal feature combination, the final model can more fully utilize the feature information in cfDNA, capturing more complex feature associations between cancer samples and normal samples, thereby improving the model's classification accuracy and generalization ability. This allows the constructed early cancer screening model to more accurately identify cancer samples in practical applications, reducing missed diagnoses and misdiagnoses, and providing more reliable technical support for early cancer screening.
[0044] In another preferred embodiment of the present invention, if the performance scores of all incremental models are less than or equal to the performance score corresponding to the initial classification model, then the initial classification model is directly used as the final model.
[0045] After performing incremental training on each feature in the common feature set and obtaining the performance score of the corresponding incremental model, the performance scores of all incremental models are compared one by one with the performance score of the initial classification model. If the performance score of any incremental model does not exceed that of the initial classification model, that is, the score of each incremental model is less than or equal to the score of the initial model, it indicates that among the currently selected common features, none of the features, or their individual addition, can improve the model performance, and may even lead to no change or a decrease in performance. In this case, based on the principle of "not performing ineffective optimization," the initial classification model is directly determined as the final model.
[0046] To ensure model performance doesn't degrade, unnecessary feature additions and model adjustments should be avoided to prevent the introduction of redundant features or overfitting due to blind optimization. The benefit is that when incremental training fails to improve model performance, unnecessary optimization steps can be stopped promptly, preserving the stability and reliability of the initial model, avoiding resource waste, and ensuring that the final model used for early cancer screening maintains at least its initial classification ability. This prevents the model from reducing its ability to distinguish between cancer and normal samples due to forcibly adding invalid features, thus ensuring that the early screening model maintains the necessary accuracy and stability in practical applications, providing reliable technical support for early cancer screening.
[0047] The present invention also includes a system for constructing a cancer early screening model based on differential statistics, for implementing the above-described method for constructing a cancer early screening model based on differential statistics, comprising: The data acquisition module is used to acquire normal sample groups and cancer sample groups, wherein each normal sample group and cancer sample group contains multiple sample data, and the corresponding original feature set of circulating cell-free DNA is extracted for each sample data. The data analysis module is used to integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples. Clustering is performed on the joint dataset to obtain several class clusters. The proportion of cancer samples in any class cluster is counted, and all class clusters are sorted according to the proportion. The top N class clusters are selected as target class clusters. The dataset of all normal samples is extracted from any target class cluster and labeled as the target sample group. N is a preset quantity threshold. The data optimization module is used to remove the target sample group from the normal sample group to obtain the comparison sample group. The target sample group and the comparison sample group are input into a preset tree model to obtain the importance score of any original feature. The original features with an importance score greater than or equal to a preset threshold are labeled as undetermined features to obtain the first undetermined feature set. The cancer sample group and the comparison sample group are input into the preset tree model, and the above process is repeated to obtain the second undetermined feature set. The model building module is used to obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and build an initial classification model based on the target feature set; The model optimization module is used to perform incremental training on each common feature in the common feature set based on the initial classification model, to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. The performance score of all incremental models is determined based on the best-in-best solution distance method, and the incremental model with the largest performance score is selected as the final model.
[0048] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A method for constructing a cancer early screening model based on differential statistics, characterized in that, Includes the following steps: S1, Obtain normal sample group and cancer sample group, wherein the normal sample group and cancer sample group each contain multiple sample data, and extract the corresponding original feature set of circulating cell-free DNA for each sample data; S2, integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples, perform clustering on the joint dataset to obtain several category clusters; count the proportion of cancer samples in any category cluster, sort all category clusters according to the proportion, select the top N category clusters as target category clusters, extract the dataset of all normal samples from any target category cluster and label it as the target sample group; N is a preset quantity threshold; S3, remove the target sample group from the normal sample group to obtain the comparison sample group, input the target sample group and the comparison sample group into the preset tree model to obtain the importance score of any original feature, and label the original features with an importance score greater than or equal to a preset threshold as undetermined features to obtain the first undetermined feature set; Input the cancer sample group and the comparison sample group into the preset tree model, repeat the above process, and obtain the second undetermined feature set; S4, obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and construct an initial classification model based on the target feature set; S5. Based on the initial classification model, perform incremental training on each common feature in the common feature set to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. Determine the performance score of all incremental models based on the best-in-best solution distance method, and select the incremental model with the largest performance score as the final model.
2. The method for constructing a cancer early screening model based on differential statistics according to claim 1, characterized in that, In S1, the original feature set includes fragment size distribution parameters, terminal base sequence fundamental frequency features, breakpoint base sequence fundamental frequency features, and nucleosome footprint features.
3. The method for constructing a cancer early screening model based on differential statistics according to claim 1, characterized in that, S2 further includes removing a target category cluster if the total number of samples in any target cluster is less than the preset total number of samples.
4. The method for constructing a cancer early screening model based on differential statistics according to claim 1, characterized in that, In S3, the specific process for obtaining important scores is as follows: The target sample group and the comparison sample group are labeled with different group labels, and together with the original feature data, they form a labeled dataset. The labeled dataset is input into a preset tree model. According to the splitting rules of the tree model, the amount of impurity reduction caused by the splitting of each original feature at the tree node is calculated. The amount of impurity reduction of the same original feature in all trees is accumulated and normalized. The data are then sorted in descending order according to the numerical value to obtain the importance score corresponding to each original feature.
5. The method for constructing a cancer early screening model based on differential statistics according to claim 1, characterized in that, The specific construction process of the initial classification model is as follows: A feature filtering operation is performed on the joint dataset, retaining only the original features in the target feature set and removing all other original features. The filtered joint dataset is then divided into a training set and a test set according to a preset ratio. A preset general model is trained based on the training set to obtain an initial classification model.
6. The method for constructing a cancer early screening model based on differential statistics according to claim 1, characterized in that, S5 further includes obtaining the performance score corresponding to the initial classification model. If the performance score of any incremental model is greater than the performance score corresponding to the initial classification model, the common feature corresponding to the incremental model is retained. If the performance score of any incremental model is less than or equal to the performance score of the initial classification model, then the common features corresponding to that incremental model are removed.
7. The method for constructing a cancer early screening model based on differential statistics according to claim 6, characterized in that, The common features retained are combined using an exhaustive method to generate several experimental groups. Incremental training is performed on each experimental group based on the initial classification model to obtain the incremental model and performance score for each experimental group. The incremental model with the highest performance score is selected as the final model.
8. The method for constructing a cancer early screening model based on differential statistics according to claim 6, characterized in that, If the performance scores of all incremental models are less than or equal to the performance score of the initial classification model, then the initial classification model is directly used as the final model.
9. A system for constructing a cancer early screening model based on differential statistics, used to implement the method for constructing a cancer early screening model based on differential statistics as described in any one of claims 1-8, characterized in that, include: The data acquisition module is used to acquire normal sample groups and cancer sample groups, wherein each normal sample group and cancer sample group contains multiple sample data, and the corresponding original feature set of circulating cell-free DNA is extracted for each sample data. The data analysis module is used to integrate the original feature sets of the normal sample group and the cancer sample group to construct a joint dataset containing the features of all samples. Clustering is performed on the joint dataset to obtain several class clusters. The proportion of cancer samples in any class cluster is counted, and all class clusters are sorted according to the proportion. The top N class clusters are selected as target class clusters. The dataset of all normal samples is extracted from any target class cluster and labeled as the target sample group. N is a preset quantity threshold. The data optimization module is used to remove the target sample group from the normal sample group to obtain the comparison sample group, input the target sample group and the comparison sample group into the preset tree model to obtain the importance score of any original feature, and label the original features with an importance score greater than or equal to a preset threshold as undetermined features to obtain the first undetermined feature set. Input the cancer sample group and the comparison sample group into the preset tree model, repeat the above process, and obtain the second undetermined feature set; The model building module is used to obtain the common feature set of the first undetermined feature set and the second undetermined feature set, remove the common feature set from the second undetermined feature set to obtain the target feature set, and build an initial classification model based on the target feature set; The model optimization module is used to perform incremental training on each common feature in the common feature set based on the initial classification model, to obtain the incremental model and performance index corresponding to each common feature. The performance index includes accuracy, AUC and loss value. The performance score of all incremental models is determined based on the best-in-best solution distance method, and the incremental model with the largest performance score is selected as the final model.
Citation Information
Patent Citations
Cancer prediction system based on integrated learning
CN115985503A
Colorectal cancer early screening method based on multi-omics sequencing
CN119274655A
MiRNA and disease relation prediction system and method based on causal feature selection
CN119905149A
Gene state cancer clustering analysis method based on deep learning
CN120220798A
Cancer detection model and construction method therefor, and reagent kit
US20240347131A1