A comprehensive analysis platform for tumor marker detection data
By designing a comprehensive analysis platform for tumor marker detection data, combining multiple omics data and clinical pathological data, a tumor prognosis model is constructed and the potential of markers is evaluated, and the problem that existing platforms cannot identify new tumor markers is solved, achieving more accurate diagnosis and personalized treatment.
Patent Information
- Application Number
- CN202510319732.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The existing comprehensive analysis platform for tumor marker detection data can only comprehensively analyze existing markers, and cannot identify new tumor markers, resulting in biased diagnostic results, insufficient prognosis evaluation and poor treatment effect.
A comprehensive analysis platform was designed, including data collection module, data coding module, model building module, potential evaluation module and marker determination module. By comprehensively analyzing a variety of omics data and clinical pathological data, a tumor prognosis model is constructed, marker potential is evaluated, and the clinical utility of candidate markers is determined.
The platform can accurately predict patients' survival and recurrence risks, identify new potential tumor markers, promote the discovery of new therapeutic targets, and improve patients' prognosis evaluation and treatment effectiveness.
Smart Images

Figure CN119851775B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedical technology, and more particularly, to a comprehensive analysis platform for tumor marker detection data. Background Art
[0002] With the progress of medical technology, the detection of tumor markers has become an important means for early diagnosis, prognosis evaluation, and treatment monitoring of tumors; tumor markers refer to biomarkers in tumor cells or patients, and their changes can reflect the presence, development, and treatment response of tumors; traditional detection methods often rely on a single marker. However, the complexity and heterogeneity of tumors make the detection effect of a single marker limited; in recent years, with the development of multi-omics technologies, such as genomics, transcriptomics, and proteomics, researchers have gradually realized the importance of comprehensive analysis by combining multiple tumor markers; therefore, how to effectively identify more tumor markers has become the key to promoting personalized medicine and precision treatment.
[0003] The patent application with the publication number CN118186086A discloses a comprehensive analysis device for head and neck tumor marker detection data; it includes a data acquisition module that uses high-throughput sequencing technology to obtain gene expression data and combines mass spectrometry or immunoassay to obtain serum marker data; a data preprocessing module that screens out genes and markers related to head and neck tumors; a feature extraction module that extracts biological process and pathway information related to head and neck tumors and obtains protein markers; a model training module that uses machine learning methods to establish a head and neck tumor marker data analysis model; a result analysis module that uses biostatistical methods to verify and evaluate the model output; this invention can efficiently and accurately analyze head and neck tumor marker detection data of patients and provide valuable clinical information.
[0004] However, although the above technology can achieve the comprehensive analysis of tumor marker detection data, it can only comprehensively analyze existing tumor markers and cannot identify new tumor markers; this leads to deviation in the diagnosis results, inability to comprehensively evaluate the tumor status, resulting in insufficient prognosis evaluation and poor treatment effect, and at the same time missing the opportunity to discover new treatment targets.
[0005] In view of this, the present invention proposes a comprehensive analysis platform for tumor marker detection data to solve the above problems. Summary of the Invention
[0006] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention provides the following technical solution: A comprehensive analysis platform for tumor marker detection data, comprising:
[0007] A data collection module for collecting marker detection data of candidate markers;
[0008] A data encoding module for encoding marker detection data to obtain marker encoded data;
[0009] A model construction module for constructing a tumor prognosis model based on the marker encoded data;
[0010] A potential evaluation module for evaluating the potential of a marker based on the marker encoded data and the tumor prognosis model;
[0011] A marker determination module for determining whether a candidate marker can be used as a tumor marker based on the marker potential. If so, perform quantitative calculation of clinical practicability.
[0012] Further, the marker detection data includes single nucleotide polymorphism detection data, clinical detection data, protein expression data, pathological detection data, and postoperative follow-up data; the clinical detection data includes laboratory detection data and intraoperative tumor data; the pathological detection data includes cytological data and histological data; the postoperative follow-up data includes the survival and recurrence of patients; the survival includes overall survival and disease-free survival;
[0013] The method for obtaining the marker encoded data includes:
[0014] Taking the single nucleotide polymorphism detection data, clinical detection data, protein expression data, and pathological detection data in the marker detection data as analysis detection data; marking the data that is not a numerical value in the analysis detection data as data to be encoded; counting the quantity corresponding to each different data in the data to be encoded and marking it as the data quantity; taking the data quantity of each different data in the data to be encoded as the corresponding data frequency; comparing all the data frequencies in turn, marking the data to be encoded whose data frequency is different from all the other data frequencies as direct data, and marking the data to be encoded with the same data frequency as combined data; taking the data frequency of the direct data as the corresponding data encoding; dividing the data corresponding to the same set of analysis detection data in the combined data into a combined set, and the combined set corresponds one-to-one with the analysis detection data; merging the data frequencies corresponding to each data in each combined set as the set encoding corresponding to the combined set; In the set of marker detection data, each combined set is replaced with the corresponding set encoding, and each direct data is replaced with the corresponding data encoding, and the replaced marker detection data is marked as marker encoded data. In the set of marker detection data, each combined set is replaced with the corresponding set encoding, and each direct data is replaced with the corresponding data encoding, and the replaced marker detection data is marked as marker encoded data.
[0015] Further, the steps for constructing the tumor prognosis model include:
[0016] Step S101: Set the number of training samples using a search optimization algorithm ;
[0017] Step S102: Randomly select a set of marker coding data as training samples;
[0018] Step S103: Construct a tumor prognosis model based on the training samples, and the tumor prognosis model is a random forest regression model;
[0019] In the said step S101, the steps of setting the number of training samples include:
[0020] Step S201: Set the search interval and the search accuracy ;
[0021] Step S202: Determine the search sequence ;
[0022] Step S203: Determine the partitioning coefficient from the search sequence ;
[0023] Step S204: Partition the search interval according to the partitioning coefficient to obtain partition points and ;
[0024] Step S205: Calculate the training accuracies corresponding to the partition points and respectively;
[0025] Step S206: Update the search interval according to the training accuracy;
[0026] Step S207: Calculate the interval width of the search interval ;
[0027] Step S208: Compare the interval width with the search accuracy . If , then . If , then jump back to step S203.
[0028] Furthermore, in the said step S202, the search sequence ;
[0029] In the said step S203, the method for determining the partitioning coefficient from the search sequence includes:
[0030] Calculate the initial partitioning coefficient according to the search interval; Take the search sequence Subtract the initial partitioning coefficient from each value in to obtain a coefficient difference; sort each coefficient difference from smallest to largest, and mark the coefficient difference at the front as the nearest difference; use the value in the search sequence corresponding to the nearest difference as the partitioning coefficient
[0031] The initial partitioning coefficient has the following expression: ; where is the maximum value of the search interval, and is the minimum value of the search interval;
[0032] In the step S204, the partitioning point ; where are the values of the first two positions in the search sequence ranked by the partitioning coefficient ;
[0033] The partitioning point ; where is the value of the first position in the search sequence ranked by the partitioning coefficient ;
[0034] Furthermore, in the step S205, the method for calculating the training accuracy corresponding to the partitioning points and includes:
[0035] Count the number of data of different types in the marker detection data and mark it as the type quantity; obtain the number of groups of marker detection data and mark it as the data quantity; use the type quantity, data quantity, and the value of the partitioning point as analysis data, and input the analysis data into the trained accuracy evaluation model to evaluate the corresponding training accuracy; the training process of the accuracy evaluation model includes:
[0036] Pre-collect groups of analysis data, set corresponding training accuracies for groups of analysis data, is an integer greater than 1, convert the analysis data and the corresponding training accuracy into a corresponding set of feature vectors; use each set of feature vectors as the input of the accuracy evaluation model, the accuracy evaluation model outputs a set of predicted training accuracies corresponding to each set of analysis data, uses the actual training accuracy corresponding to each set of analysis data as the prediction target, and the actual training accuracy is the pre-set training accuracy corresponding to the analysis data; use minimizing the sum of prediction errors of all analysis data as the training target; train the accuracy evaluation model until the sum of prediction errors reaches convergence and then stop training; the accuracy evaluation model is a deep neural network model;
[0037] In step S206, the method for updating the search interval according to the training accuracy includes:
[0038] Mark the training accuracy corresponding to the division point as the first accuracy , and mark the training accuracy corresponding to the division point as the second accuracy ; Compare the first accuracy with the second accuracy ; If , update the minimum value of the search interval to , and do not update the maximum value; If , update the maximum value of the search interval to , and do not update the minimum value;
[0039] In step S207, the method for calculating the interval width of the search interval is: subtract the minimum value corresponding to the search interval from the maximum value to obtain the interval width of the search interval .
[0040] Furthermore, in step S103, the method for constructing a tumor prognosis model according to training samples includes:
[0041] Initialize the model parameters, where the model parameters include the maximum depth of the decision tree and the minimum number of samples in the internal nodes; Delete the postoperative follow-up data in each group of training samples and mark it as training data; Set different digital labels for the postoperative follow-up data of different groups and mark it as postoperative labels; Sample tree subsets from groups of training data, and each tree subset includes groups of training data, where , ; Train decision trees respectively according to the tree subsets, where the training process of each decision tree is independent and the same, and finally construct a random forest regression model containing
[0042] Input each group of training data into each decision tree in turn, record the decision output, and the decision output is the postoperative label output by each decision tree, and mark it as the prediction label; Calculate the mean value of the prediction labels corresponding to the same training data and use it as the predicted value; Calculate the mean square error between each predicted value and the true value and use it as the loss function value, where the true value is the postoperative label corresponding to the training data input into the decision tree; Use the random search method or the Bayesian optimization method to perform model hyperparameter tuning, and select the model hyperparameters with the smallest loss function value as the model hyperparameters of the tumor prognosis model.
[0043] Further, the method of training a decision tree using a subset of trees includes:
[0044] Dividing the group of training data in a subset of trees into a training set and a validation set , marking the training data in the training set as , , where is the number of groups of training data in the training set ; using the feature to represent the data in the training data, where is the number of different types of data in the training data; calculating the information gain ratio for each feature
[0045] Selecting the feature corresponding to the maximum information gain ratio as the internal node, and dividing the training set into decision sub-sets according to the feature corresponding to the maximum information gain ratio ; for each decision sub-set divided according to the maximum information gain ratio , repeating the algorithm recursion process until the number of training data in the decision sub-set is less than or equal to the minimum number of samples of the internal node, or the number of recursions is greater than or equal to the maximum depth of the decision tree, at which point the recursion ends; evaluating the trained decision tree model using the validation set to finally complete the training of the decision tree.
[0046] Further, the expression for the information gain ratio is: ; where is the decision sub-set obtained by dividing the training set according to the feature , , is the number of decision sub-sets, is 's information entropy, is the information entropy of the training set ;
[0047] The expression for is:
[0048] The expression for is: represents the th Postoperative labels corresponding to the group of training data.
[0049] Further, the method for evaluating the potential of the biomarker includes:
[0050] Taking the biomarker coding data not labeled as training samples as evaluation samples, and labeling the postoperative follow-up data in the evaluation samples as real data; deleting the postoperative follow-up data in each evaluation sample and labeling it as evaluation data; inputting each group of evaluation data into the trained tumor prognosis model respectively to predict the corresponding postoperative labels; obtaining the corresponding postoperative follow-up data according to the predicted postoperative labels and labeling it as predicted data; taking each real data and the corresponding predicted data as a set of analysis sets, and the analysis sets correspond one-to-one with the real data; subtracting each value in each real data from the corresponding value in the corresponding predicted data respectively and taking the absolute value to obtain the value difference.
[0051] Presetting a weight set, where the weight set includes the weight coefficients corresponding to each data in the postoperative follow-up data; multiplying each value difference by the corresponding weight coefficient to obtain the weighted difference; adding up the weighted differences corresponding to each group of analysis sets in sequence to obtain the total value difference; counting the number of analysis sets and labeling it as the set number; adding up each total value difference in sequence and then dividing by the set number to obtain the mean of the total differences, and taking the reciprocal of the mean of the total differences as the biomarker potential.
[0052] Further, the method for determining whether a candidate biomarker can be used as a tumor biomarker includes:
[0053] Presetting a potential threshold, and comparing the biomarker potential with the potential threshold; if the biomarker potential is greater than or equal to the potential threshold, determining that the candidate biomarker is a tumor biomarker; if the biomarker potential is less than the potential threshold, determining that the candidate biomarker is not a tumor biomarker.
[0054] The method for performing quantitative calculation of clinical utility includes:
[0055] Calculating the correlation between the protein expression data in a group of biomarker detection data and each data in the postoperative follow-up data; multiplying each correlation by the corresponding weight coefficient to obtain the weighted degree of correlation; taking the absolute value of each weighted degree of correlation and adding them up in sequence to obtain the clinical utility corresponding to the candidate biomarker.
[0056] The expression formula for the correlation is: ; where is the correlation, is the protein expression data in the th group of biomarker detection data, is the th data in the postoperative follow-up data in the th group of biomarker detection data.
[0057] Technical effects and advantages of a comprehensive analysis platform for tumor marker detection data according to the present invention:
[0058] By comprehensively analyzing various omics data such as genomics, transcriptomics, and proteomics, clinical pathological data, and postoperative follow-up data, and using optimized machine learning algorithms to construct a tumor prognosis model, it is possible to accurately predict the survival status and recurrence risk of patients, precisely evaluate the potential of markers, thereby effectively identifying new potential tumor markers, promoting the discovery of new treatment targets, and further promoting the realization of early diagnosis and personalized treatment; at the same time, using a correlation algorithm to objectively quantify the clinical value of candidate markers provides a basis for selecting the most suitable tumor markers for clinical applications; not only improves the detection and verification efficiency of tumor markers, but also provides important support for personalized medicine and precision treatment, and can significantly improve the prognosis evaluation and treatment effect of patients. Brief Description of the Drawings
[0059] Figure 1 It is a schematic diagram of a comprehensive analysis platform for tumor marker detection data according to Embodiment 1 of the present invention;
[0060] Figure 2 It is a flowchart of a method for constructing a tumor prognosis model according to Embodiment 1 of the present invention. Detailed Embodiments
[0061] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0062] Embodiment 1
[0063] Please refer to Figure 1 As shown, a comprehensive analysis platform for tumor marker detection data described in this embodiment includes a data collection module, a data encoding module, a model construction module, a potential evaluation module, and a marker determination module; each module is connected by wired and / or wireless means to achieve data transmission between modules.
[0064] The data collection module is used to collect marker detection data of candidate markers.
[0065] Candidate biomarkers are biomarkers that are identified as potentially tumor-related in a study but have not been confirmed for clinical use, such as HLA-DRA, HLA-DRB1, etc.; biomarker detection data include single nucleotide polymorphism detection data, clinical detection data, protein expression data, pathological detection data, and postoperative follow-up data; a total of groups are integers greater than 1.
[0066] Single nucleotide polymorphism detection data provide genetic variation information of gene loci, such as single nucleotide polymorphism identifiers, single nucleotide polymorphism positions, allele information, etc., and are obtained through whole genome sequencing or high-throughput SNP chip technology; single nucleotide polymorphism detection data are used to identify gene variations that affect tumor susceptibility or tumor progression, and help screen tumor-related biomarkers.
[0067] Clinical detection data include laboratory test data and intraoperative tumor data; laboratory test data such as blood test data (such as blood cell count, biochemical indicators, etc.), imaging test data, etc., imaging test data such as CT scan data (such as tumor size, tumor shape, etc.), MRI data (such as the boundary and infiltration of the tumor), ultrasound examination data (such as the echo characteristics of the tumor, blood flow dynamics, etc.), etc.; laboratory test data are obtained through the hospital laboratory; intraoperative tumor data such as tumor location, tumor size, tumor stage, etc., are obtained through the observation of doctors during tumor surgery; clinical detection data are used to comprehensively analyze tumor characteristics and provide an important basis for the subsequent clinical validation of candidate biomarkers.
[0068] Protein expression data are the protein expression levels of candidate biomarkers and are obtained through immunohistochemistry or Western blotting; protein expression data are used to verify whether it is upregulated or downregulated in tumor tissues and further understand the role of candidate biomarkers in tumor development.
[0069] Pathological detection data include cytological data and histological data; cytological data such as cell size, cell shape, etc., are obtained through flow cytometry or microscopic imaging; histological data such as tumor type, tumor grade, tumor stage, etc., are obtained through tissue section or immunohistochemistry; pathological detection data are used to evaluate the expression of candidate biomarkers in tumor tissues and help verify their potential as tumor biomarkers.
[0070] Postoperative follow-up data include the survival and recurrence status of patients, which are obtained through telephone interviews or outpatient follow-up; the survival status includes overall survival and disease-free survival. Overall survival is the time span from the start of treatment to the death of the patient, and disease-free survival is the time span from the start of treatment to the first progression of the disease (such as tumor enlargement or the emergence of new conditions); the recurrence status is the time span when the tumor reappears after treatment; postoperative follow-up data are used to judge the treatment effect of tumors and the long-term survival rate of patients, and to verify the correlation between candidate markers and the prognosis of patients. For example, certain markers may be closely related to the survival period and recurrence rate of patients.
[0071] The data encoding module is used to encode the marker detection data to obtain marker encoding data.
[0072] The methods for obtaining marker encoding data include:
[0073] Take the single nucleotide polymorphism detection data, clinical detection data, protein expression data, and pathological detection data in the marker detection data as analysis detection data; in the analysis detection data, data that is not a numerical value is marked as data to be encoded; count the quantity corresponding to each different data in the data to be encoded and mark it as the data quantity; for example, the quantity corresponding to the tumor shape being round, the quantity corresponding to the tumor shape being multinodular, etc.; take the data quantity of each different data in the data to be encoded as the corresponding data frequency; compare all the data frequencies in sequence, mark the data to be encoded whose data frequency is different from all the other data frequencies as direct data, and mark the data to be encoded with the same data frequency as combined data; take the data frequency of the direct data as the corresponding data encoding; in the combined data, the data corresponding to the same set of analysis detection data are divided into a combined set, and the combined set corresponds one-to-one with the analysis detection data; merge the data frequencies corresponding to each data in each combined set as the set encoding corresponding to the combined set; in In a set of marker detection data, each combined set is replaced with the corresponding set encoding, and each direct data is replaced with the corresponding data encoding. Mark the replaced marker detection data as marker encoding data.
[0074] Exemplarily, there are combined set A and combined set B. Combined set A includes tumor shape A and single nucleotide polymorphism position B, and combined set B includes tumor shape B, tumor type C, and tumor grade D; the data frequencies corresponding to tumor shape A and tumor shape B are 10, and the data frequencies corresponding to single nucleotide polymorphism position B, tumor type C, and tumor grade D are 20; therefore, the set encoding corresponding to combined set A is 1020, and the set encoding corresponding to combined set B is 102020.
[0075] The model construction module is used to construct a tumor prognosis model based on the marker encoding data.
[0076] As shown Figure 2 below, the steps of constructing a tumor prognosis model include:
[0077] Step S101: Set the number of training samples using a search optimization algorithm ;
[0078] Step S102: Randomly select groups of marker coding data as training samples;
[0079] Step S103: Construct a tumor prognosis model based on the training samples, and the tumor prognosis model is a random forest regression model.
[0080] In the above step S101, the steps of setting the number of training samples include:
[0081] Step S201: Set the search interval and the search accuracy ;
[0082] Step S202: Determine the search sequence ;
[0083] Step S203: Determine the partition coefficient from the search sequence ;
[0084] Step S204: Divide the search interval according to the partition coefficient to obtain the partition points and ;
[0085] Step S205: Calculate the training accuracies corresponding to the partition points and respectively;
[0086] Step S206: Update the search interval according to the training accuracy; Step S207: Calculate the interval width of the search interval ;
[0087] Step S208: Compare the interval width with the search accuracy . If , then . If , then jump back to step S203.
[0088] In the above step S201, the search accuracy is preset by those skilled in the art according to the actual situation.
[0089] In the above step S202, search for the sequence .
[0090] In the above step S203, the method for determining the partitioning coefficient from the search sequence includes:
[0091] Calculate the initial partitioning coefficient according to the search interval; subtract the initial partitioning coefficient from each value in the search sequence to obtain the coefficient difference; sort each coefficient difference from smallest to largest, and mark the coefficient difference at the front as the nearest difference; use the value in the search sequence corresponding to the nearest difference as the partitioning coefficient ;
[0092] The expression of the initial partitioning coefficient is: ; where is the maximum value of the search interval, and is the minimum value of the search interval.
[0093] In the above step S204, the partitioning point ; where are the values of the first two positions in the search sequence ranked before the partitioning coefficient ;
[0094] The partitioning point ; where is the value of the position before the partitioning coefficient in the search sequence.
[0095] In the above step S205, the method for calculating the training accuracies corresponding to the partitioning points and includes:
[0096] Count the number of data of different types in the marker detection data and mark it as the type quantity; obtain the number of groups of the marker detection data and mark it as the data group quantity; use the type quantity, the data group quantity, and the value of the partitioning point as the analysis data, and input the analysis data into the trained accuracy evaluation model to evaluate the corresponding training accuracy; the training process of the accuracy evaluation model includes:
[0097] Pre-collect groups of analysis data, set the corresponding training accuracy for groups of analysis data, is an integer greater than 1, convert the analysis data and the corresponding training accuracy into a corresponding set of feature vectors; the training accuracy corresponding to the analysis data is collected by those skilled in the art during the process of historically training the random forest regression model, Group the analysis data. According to the values of the division points in each group of analysis data, select the corresponding number of marker coding data as training samples to train the random forest regression model. After the model training is completed, evaluate the model accuracy of each random forest regression model in turn according to the actual situation, and use it as the training accuracy of the corresponding analysis data. For each group of analysis data, set the corresponding training accuracy in turn;
[0098] Take each group of feature vectors as the input of the accuracy evaluation model. The accuracy evaluation model takes a set of predicted training accuracies corresponding to each group of analysis data as the output, takes the actual training accuracy corresponding to each group of analysis data as the prediction target, and the actual training accuracy is the pre-set training accuracy corresponding to the analysis data; take minimizing the sum of the prediction errors of all analysis data as the training target; where the calculation formula of the prediction error is , where is the prediction error, is the group number of the feature vector corresponding to the analysis data, is the predicted training accuracy corresponding to the \(i\)-th group of analysis data, is the actual training accuracy corresponding to the \(i\)-th group of analysis data; train the accuracy evaluation model until the sum of the prediction errors reaches convergence and then stop training.
[0099] The above accuracy evaluation model is specifically a deep neural network model; it includes an input layer, a hidden layer, and an output layer; each hidden layer includes multiple neurons, and there are connections between each neuron and the neurons in the next layer. The connections contain weights that determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer. The activation function introduces non-linearity and allows the network to learn more complex patterns and features.
[0100] In the above step S206, the method for updating the search interval according to the training accuracy includes:
[0101] Mark the training accuracy corresponding to the division point as the first accuracy , and mark the training accuracy corresponding to the division point as the second accuracy ; compare the first accuracy and the second accuracy ; if , update the minimum value of the search interval to , and do not update the maximum value; if , update the maximum value of the search interval to , and do not update the minimum value; Exemplarily, the search interval is , if , the search interval is updated to , if , the search interval is updated to .
[0102] In the above step S207, the method for calculating the interval width of the search interval is: subtract the minimum value corresponding to the maximum value of the search interval to obtain the interval width of the search interval .
[0103] In the above step S103, the method for constructing a tumor prognosis model based on training samples includes:
[0104] Initialize the model parameters, where the model parameters include the maximum depth of the decision tree and the minimum number of samples in the internal nodes; the model parameters are all preset by those skilled in the art according to the actual situation; delete the postoperative follow-up data in each group of training samples and mark them as training data; from groups of training data, sample tree subsets, each tree subset contains groups of training data, and train decision trees respectively according to tree subsets; among them, , .
[0105] The method for training a decision tree with a tree subset includes:
[0106] Divide the groups of training data in a tree subset into a training set and a validation set , mark the training data in the training set as , , is the number of groups of training data in the training set ; use the feature to represent the data in the training data, , is the number of different types of data in the training data; calculate the information gain rate of each feature ;
[0107] The information gain rate is expressed as: ; in the formula, is the decision subset divided by the training set according to the feature , , is the number of decision subsets, is The information entropy of is the information entropy of the training set ;
[0108] The expression of is:
[0109] The expression of is: In the formula, represents the postoperative label corresponding to the th group of training data in the decision subset . The postoperative label is the digital label corresponding to the postoperative follow-up data, and the digital labels corresponding to the postoperative follow-up data of different groups are all different;
[0110] Select the feature corresponding to the maximum information gain rate as the internal node, and divide the training set into decision subsets according to the feature corresponding to the maximum information gain rate ; For each decision subset divided according to the maximum information gain rate , repeat the recursive process of the algorithm until the number of training data in the decision subset is less than or equal to the minimum number of samples of the internal node, or the number of recursions is greater than or equal to the maximum depth of the decision tree, and the recursion ends; Evaluate the trained decision tree model using the validation set to finally complete the training of the decision tree.
[0111] Use the remaining tree subsets in turn to train decision trees, where the training process of each decision tree is independent and the same, and finally construct a random forest regression model containing decision trees.
[0112] Input each group of training data into each decision tree in turn, record the decision output, and the decision output is the postoperative label output by each decision tree, and mark it as the predicted label; Calculate the mean value of the predicted labels corresponding to the same training data and use it as the predicted value; Calculate the mean square error between each predicted value and the true value and use it as the loss function value, and the true value is the postoperative label corresponding to the training data input into the decision tree; Use the random search method or the Bayesian optimization method to optimize the model hyperparameters, reduce the loss function value, and select the model hyperparameters with the smallest loss function value as the model hyperparameters of the tumor prognosis model; It should be noted that both the random search method and the Bayesian optimization method are existing technologies and will not be elaborated here.
[0113] Potential evaluation module, used to evaluate the biomarker potential according to the biomarker coding data and the tumor prognosis model.
[0114] Methods for evaluating the potential of biomarkers include:
[0115] Taking the biomarker coding data not labeled as training samples as evaluation samples, and labeling the postoperative follow-up data in the evaluation samples as real data; deleting the postoperative follow-up data in each evaluation sample and labeling it as evaluation data; inputting each group of evaluation data into the trained tumor prognosis model respectively to predict the corresponding postoperative labels; obtaining the corresponding postoperative follow-up data according to the predicted postoperative labels and labeling it as predicted data; taking each real data and the corresponding predicted data as a group of analysis sets, and the analysis sets correspond one-to-one with the real data; subtracting each value in each real data from the corresponding value in the corresponding predicted data respectively and taking the absolute value to obtain the value difference;
[0116] Presetting a weight set, where the weight set includes the weight coefficients corresponding to each data in the postoperative follow-up data, and the weight set is preset by those skilled in the art according to the actual situation; multiplying each value difference by the corresponding weight coefficient to obtain the weighted difference; adding up the weighted differences corresponding to each group of analysis sets in sequence to obtain the total value difference; counting the number of analysis sets and labeling it as the set number; adding up each total value difference in sequence and then dividing by the set number to obtain the mean value of the total difference, and taking the reciprocal of the mean value of the total difference as the biomarker potential.
[0117] A biomarker determination module, which is used to determine whether a candidate biomarker can be used as a tumor biomarker according to the biomarker potential. If so, a quantitative calculation of clinical practicability is performed.
[0118] Methods for determining whether a candidate biomarker can be used as a tumor biomarker include:
[0119] Presetting a potential threshold, which is preset by those skilled in the art according to the actual situation; comparing the biomarker potential with the potential threshold; if the biomarker potential is greater than or equal to the potential threshold, determining that the candidate biomarker is a tumor biomarker; if the biomarker potential is less than the potential threshold, determining that the candidate biomarker is not a tumor biomarker.
[0120] Methods for performing a quantitative calculation of clinical practicability include:
[0121] Calculating the correlation between the protein expression data in a group of biomarker detection data and each data in the postoperative follow-up data; multiplying each correlation by the corresponding weight coefficient to obtain the weighted degree of correlation; taking the absolute value of each weighted degree of correlation respectively and adding them up in sequence to obtain the clinical practicability corresponding to the candidate biomarker;
[0122] The expression formula for the correlation is: ; where is the correlation, is the protein expression data among the detection data of the first group of biomarker detection data, and is a data within the postoperative follow-up data among the detection data of the first group of biomarker detection data.
[0123] In this embodiment, by comprehensively analyzing various omics data such as genomics, transcriptomics, and proteomics, clinical pathological data, and postoperative follow-up data, and using an optimized machine learning algorithm to construct a tumor prognosis model, it is possible to accurately predict the survival situation and recurrence risk of patients, precisely evaluate the potential of biomarkers, thereby effectively identifying new potential tumor biomarkers, promoting the discovery of new treatment targets, and further promoting the realization of early diagnosis and personalized treatment; at the same time, using a correlation algorithm to objectively quantify the clinical value of candidate biomarkers provides a basis for selecting the most suitable tumor biomarker for clinical application; not only improves the detection and verification efficiency of tumor biomarkers, but also provides important support for personalized medicine and precision treatment, and can significantly improve the prognosis evaluation and treatment effect of patients.
[0124] Example 2
[0125] The present application also provides an electronic device. The electronic device may include one or more processors and one or more memories. Among them, the memory stores computer-readable code, and when the computer-readable code is run by one or more processors, it can execute a comprehensive analysis platform for tumor biomarker detection data as described above.
[0126] The method or system according to the embodiment of the present application can also be implemented by means of the architecture of the electronic device shown in the present application. The electronic device may include a bus, one or more CPUs, a ROM, a RAM, a communication port connected to a network, an input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or a hard disk, can store a comprehensive analysis platform for tumor biomarker detection data provided by the present application. Further, the electronic device may further include a user interface. Of course, the architecture shown in the present application is only exemplary, and when implementing different devices, one or more components shown in the electronic device of the present application can be omitted according to actual needs.
[0127] Example 3
[0128] One embodiment of the present application discloses a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are run by a processor, a comprehensive analysis platform for tumor marker detection data according to the embodiments of the present application described with reference to the above drawings can be executed. The storage medium includes but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0129] In addition, according to the embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application, such as: a comprehensive analysis platform for tumor marker detection data. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0130] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0131] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A comprehensive analysis platform for tumor marker detection data, characterized in that: include: A data collection module, used to collect marker detection data of candidate markers; The marker detection data includes single nucleotide polymorphism detection data, clinical detection data, protein expression data, pathological detection data and postoperative follow-up data; A data encoding module, used to encode the marker detection data and obtain the marker encoding data; A model building module is used to build a tumor prognosis model based on marker coding data; The step of constructing a tumor prognosis model comprises: Step S101: using a search optimization algorithm to set the number of training samples b; Step S102: randomly selecting group b marker encoding data as training samples; Step S103: constructing a tumor prognosis model according to the training samples, where the tumor prognosis model is a random forest regression model; A potential assessment module is used to assess the potential of a marker based on the marker coding data and the tumor prognosis model; the method for assessing the potential of a marker includes: Different digital labels are set for different groups of postoperative follow-up data, and they are marked as postoperative labels; the marker coding data that are not marked as training samples are used as evaluation samples, and the postoperative follow-up data in the evaluation samples are marked as real data; the postoperative follow-up data in each evaluation sample is deleted and marked as evaluation data; each group of evaluation data is input into the trained tumor prognosis model to predict the corresponding postoperative label; according to the predicted postoperative label, the corresponding postoperative follow-up data is obtained and marked as predicted data; each real data and the corresponding predicted data are used as a set of analysis sets, and the analysis sets correspond to the real data one by one; each value in each real data is subtracted from the corresponding value in the corresponding predicted data, and the absolute value is taken to obtain the numerical difference; Preset a weight set, which includes a weight coefficient corresponding to each data in the postoperative follow-up data; multiply each numerical difference by the corresponding weight coefficient to obtain a weight difference; add the weight differences corresponding to each analysis set in sequence to obtain a total numerical difference; count the number of analysis sets and mark it as the number of sets; add each total numerical difference in sequence, divide it by the number of sets, obtain the mean of the total difference, and use the reciprocal of the mean of the total difference as the marker potential; The marker determination module is used to determine whether the candidate marker can be used as a tumor marker based on the marker potential, and if so, to perform quantitative calculations of clinical practicality.
2. A comprehensive analysis platform for tumor marker detection data according to claim 1, characterized in that: The clinical test data include laboratory test data and intraoperative tumor data; the pathological test data include cytological data and histological data; the postoperative follow-up data include the patient's survival and recurrence; the survival status includes overall survival and disease-free survival; The method for obtaining marker coding data comprises: The single nucleotide polymorphism detection data, clinical detection data, protein expression data and pathological detection data in the marker detection data are used as analysis detection data; the data that is not a numerical value in the analysis detection data is marked as data to be encoded; the number of each different data in the data to be encoded is counted and marked as the data number; the number of data of each different data in the data to be encoded is used as the corresponding data frequency; all data frequencies are compared in sequence, and the data to be encoded whose data frequency is different from all other data frequencies is marked as direct data, and the data to be encoded with the same data frequency is marked as combined data; the data frequency of the direct data is used as the corresponding data code; the data corresponding to the same group of analysis detection data in the combined data is divided into a combined set, and the combined set corresponds to the analysis detection data one by one; the data frequency corresponding to each data in each combined set is merged as the set code corresponding to the corresponding combined set; in group a of marker detection data, each combined set is replaced with the corresponding set code, and each direct data is replaced with the corresponding data code, and the replaced marker detection data is marked as marker coding data.
3. A comprehensive analysis platform for tumor marker detection data according to claim 2, characterized in that: In step S101, the step of setting the number of training samples b includes: Step S201: setting the search interval [1, a] and the search precision θ; Step S202: Determine the search sequence F(c); Step S203: Determine the partition coefficient d from the search sequence F(c); Step S204: Divide the search interval according to the division coefficient d to obtain division points ψ1 and ψ2; Step S205: Calculate the training accuracy corresponding to the division points ψ1 and ψ2 respectively; Step S206: updating the search interval according to the training accuracy; Step S207: Calculate the interval width g of the search interval; Step S208: Compare the interval width g with the search accuracy θ. If g≤θ, then If g>θ, jump back to step S203.
4. A comprehensive analysis platform for tumor marker detection data according to claim 3, characterized in that: In step S202, the search sequence The method of determining the partition coefficient d from the search sequence F(c) in step S203 includes: Calculate the initial partition coefficient based on the search interval Subtract the initial partition coefficient from each value in the search sequence F(c) to obtain the coefficient difference; sort each coefficient difference from small to large, and mark the coefficient difference at the front as the most recent difference; use the value in the search sequence F(c) corresponding to the most recent difference as the partition coefficient d; Initial partition coefficient The expression is: In the formula, F max is the maximum value of the search interval, F min is the minimum value of the search interval; In step S204, the division point Where d(2) is the value of the first two digits of the partition coefficient d in the search sequence; Divide point Wherein, d(1) is the value before the partition coefficient d in the search sequence.
5. A comprehensive analysis platform for tumor marker detection data according to claim 4, characterized in that: In step S205, the method for calculating the training accuracy corresponding to the division points ψ1 and ψ2 includes: The number of different types of data in the marker detection data is counted and marked as the number of types; the number of groups of the marker detection data is obtained and marked as the number of data groups; the number of types, the number of data groups and the numerical values of the dividing points are used as analysis data, and the analysis data is input into the trained accuracy assessment model to evaluate the corresponding training accuracy; the training process of the accuracy assessment model includes: Collect h groups of analysis data in advance, set corresponding training precisions for the h groups of analysis data, where h is an integer greater than 1, and convert the analysis data and the corresponding training precisions into a corresponding set of feature vectors; use each set of feature vectors as input to an accuracy assessment model, which outputs a set of predicted training precisions corresponding to each set of analysis data, and uses the actual training precision corresponding to each set of analysis data as a prediction target, where the actual training precision is the pre-set training precision corresponding to the analysis data; minimize the sum of prediction errors of all analysis data as a training target; train the accuracy assessment model until the sum of prediction errors reaches convergence and stops training; the accuracy assessment model is a deep neural network model; In step S206, the method for updating the search interval according to the training accuracy includes: The training accuracy corresponding to the division point ψ1 is marked as the first accuracy ω1, and the training accuracy corresponding to the division point ψ2 is marked as the second accuracy ω2; the first accuracy ω1 is compared with the second accuracy ω2; if ω1<ω2, the minimum value of the search interval is updated to ψ1, and the maximum value is not updated; if ω1≥ω2, the maximum value of the search interval is updated to ψ2, and the minimum value is not updated; In step S207, the method for calculating the interval width g of the search interval is: subtracting the corresponding minimum value from the maximum value of the search interval to obtain the interval width g of the search interval.
6. A comprehensive analysis platform for tumor marker detection data according to claim 5, characterized in that: In step S103, the method for constructing a tumor prognosis model according to training samples includes: Initialize model parameters, including the maximum depth of the decision tree and the minimum number of samples of internal nodes; delete the postoperative follow-up data in each group of training samples and mark them as training data; sample u tree subsets from b groups of training data, each tree subset includes l groups of training data, where b>u>1, b=u×l; train u decision trees based on u tree subsets, where the training process of each decision tree is independent and the same, and finally construct a random forest regression model containing u decision trees; Input each set of training data into each decision tree in turn, record the decision output, which is the postoperative label output by each decision tree and marked as the predicted label; calculate the mean of the predicted labels corresponding to the same training data and use it as the predicted value; calculate the mean square error between each predicted value and the true value and use it as the loss function value, where the true value is the postoperative label corresponding to the training data input into the decision tree; use random search method or Bayesian optimization method to tune the model hyperparameters, and select the model hyperparameters with the smallest loss function value as the model hyperparameters of the tumor prognosis model.
7. A comprehensive analysis platform for tumor marker detection data according to claim 6, characterized in that: Methods for training a decision tree using a subset of trees include: Divide the l sets of training data in a tree subset into training set α and validation set β, and mark the training data in training set α as α p , is the number of training data sets in the training set α; using feature k v Represents the data in the training data, v∈[1,y], y is the number of different types of data in the training data; calculate each feature k v The information gain rate z(α,k v ); Select the maximum information gain rate z(α,k v ) corresponds to the feature k v As an internal node, and according to the maximum information gain rate z(α,k v ) corresponds to the feature k v The training set α is divided into decision subsets; for each decision subset according to the maximum information gain rate z(α, k v ) and repeat the algorithm recursive process until the number of training data in the decision subset is less than or equal to the minimum number of samples of the internal nodes, or the number of recursions is greater than or equal to the maximum depth of the decision tree, and the recursion ends; the trained decision tree model is evaluated using the validation set β, and finally the training of the decision tree is completed.
8. A comprehensive analysis platform for tumor marker detection data according to claim 7, characterized in that: Information gain rate z(α,k v ) is: In the formula, αr is the training set α according to feature k v The decision subsets divided, r∈[1,R], R is the number of decision subsets, H(ar) is the information entropy of αr, and H(α) is the information entropy of the training set α; The expression of H(α) is: The expression of H(αr) is: In the formula, αr p Represents the postoperative label corresponding to the p-th group of training data in the decision subset αr.
9. A comprehensive analysis platform for tumor marker detection data according to claim 8, characterized in that: The method for determining whether a candidate marker can be used as a tumor marker comprises: A potential threshold is preset, and the potential of the marker is compared with the potential threshold; if the potential of the marker is greater than or equal to the potential threshold, the candidate marker is determined to be a tumor marker; if the potential of the marker is less than the potential threshold, the candidate marker is determined not to be a tumor marker; The method for quantitative calculation of clinical practicality includes: Calculate the correlation between the protein expression data in the marker detection data of group a and each data in the postoperative follow-up data; multiply each correlation by the corresponding weight coefficient to obtain the weight correlation degree; take the absolute value of each weight correlation degree and add them up in sequence to obtain the clinical practicality corresponding to the candidate marker; The expression of correlation is: In the formula, q is the correlation, m is i is the protein expression data in the i-th group of marker detection data, n i It is a data in the postoperative follow-up data in the i-th group of marker detection data.
Citation Information
Patent Citations
Comprehensive analysis device for head and neck tumor marker detection data
CN118186086A