Tc disease risk prediction modeling based on cell proliferation markers

By setting TI-RADS classification and TK1 test results as independent variables in physical examination samples, a TC disease risk prediction model was established, which solves the problems of difficult sample acquisition and model iteration and upgrading in existing technologies, and achieves efficient and accurate risk assessment for healthy people.

CN115705928BActive Publication Date: 2026-03-03SHENZHEN HUARUI TONGKANG BIOTECHNOLOGICAL
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-05
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

Existing risk assessment systems for thyroid cancer suffer from high costs and difficulties in obtaining samples, as well as low sample accumulation efficiency. This makes it difficult to iterate and upgrade risk assessment models, and existing methods are not sensitive or specific enough in the early detection of lesions.

Method used

A TC disease risk prediction modeling method based on cell proliferation markers was adopted. By setting the TI-RADS classification in physical examination samples as the endpoint event, and combining multiple physical examination indicators, mainly serum thymidine kinase 1 (TK1) detection results, regression modeling was used to optimize the model for high-frequency and high-throughput screening of healthy people.

Benefits of technology

It reduced the cost and difficulty of obtaining positive samples, improved the feasibility of iterative upgrades of risk prediction models, achieved universal applicability and efficient screening for healthy populations, and improved the accuracy and applicability of the models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115705928B_ABST
    Figure CN115705928B_ABST
Patent Text Reader

Abstract

The application relates to a TC disease risk prediction modeling method based on a cell proliferation marker, comprising the following steps: collecting sample data; converting a medical imaging examination result into a quantitative TI-RADS classification, and setting a TI-RADS classification of TR-3 and above as an endpoint event; taking multiple physical examination index items with a cell proliferation marker serum thymidine kinase 1 (TK1) detection result as independent variables; and modeling through a regression method, so as to obtain a thyroid cancer (TC) risk prediction model. The application sets the comprehensive classification of the imaging examination result with a determined TC risk as a risk prediction endpoint event, thereby reducing the acquisition cost and difficulty of positive result samples, and making it possible to stably accumulate positive samples. In the same time period, the iteration upgrading feasibility of the risk prediction model obtained by the system is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of disease risk prediction modeling, and in particular to a method for predicting the disease risk of TC based on cell proliferation markers. Background Technology

[0002] Thyroid cancer (TC) is one of the ten most common cancers. Its causes are complex, and early-stage TC often presents with no obvious symptoms. According to 2021 statistics, there were 201,000 thyroid cancer patients in my country in 2015. It is estimated that in 2020, there were 53,389 male and 167,704 female thyroid cancer patients in my country, accounting for 38.89% and 37.36% of the world's thyroid cancer patients, respectively. The number of new cases in 2020 increased by approximately 20,000 compared to 2015, showing a trend towards younger patients and a greater prevalence in women. Furthermore, the incidence rate in both cases is higher than the global average. From 1980 to 2015, the incidence rate of thyroid cancer in my country continued to rise, partly due to the increased pace and stress of modern life and the stronger influence of environmental factors, and partly due to the continuous upgrading of medical testing technology.

[0003] To achieve early screening for papillary thyroid carcinoma (TC), follicular TC, undifferentiated TC originating from follicular epithelial cells, and myeloid TC primarily derived from parafollicular cells, in addition to routine surgical examination, blood tests for thyroid hormones, thyroid autoantibodies, and tumor markers are necessary. For high-risk individuals, fine-needle aspiration cytology (FNA) is required to confirm tumor type and differentiation / stage. However, thyroid surgical palpation is highly dependent on physician experience and subjectivity, ultrasound only detects tissue structure externally, and the sensitivity and specificity of blood marker tests are insufficient for early screening. Furthermore, the overuse of FNA can cause unnecessary suffering for many patients. Therefore, developing a non-invasive and sustainable method for thyroid cancer risk assessment based on circulatory system tumor risk factors to reduce the need for unnecessary pathological examinations is of significant practical importance.

[0004] In recent years, thyroid cancer risk prediction models can be summarized into three main modes: 1) Based on the "signs of medical imaging (especially ultrasound technology)," such as composition, echo, morphology, boundary, calcification, etc., a receiver operating characteristic curve (ROC) is constructed to predict and assess the risk of thyroid cancer. For example, see "Chen Junhui, Zhang Man, Liu Shuipeng, et al. Study on thyroid cancer risk prediction based on multivariate logistic regression β-value integral method of ultrasound signs [J]. Chinese Journal of Cancer, 2019, 29(004): 289-293."; 2) Based on "Hashimoto's thyroiditis, thyroid function, and serum thyroid autoantibodies," the risk of thyroid cancer is assessed. For example, see "Kadam S, Ghadhban B, Abduljabbarsaleh M. Hashimoto's Thyroiditis Increases Risk for Differentiated Thyroid Carcinoma [J]. Annals of the Romanian Society for Cell Biology, 2021, 25: 5952-5961.; 3) Based on RNAseq data from cancer genome mapping databases such as TGGA and KEGG pathway enrichment analysis, construct a risk assessment and artificial neural network model for thyroid cancer based on variant pathways. See references such as "Zhao Y, Zhao L, Mao T, et al. Assessment of risk based on variant pathways and establishment of an artificial neural network model of thyroid cancer[J]. BMC Medical Genetics, 2019, 20.".

[0005] Regarding the aforementioned technologies, the inventors believe they suffer from the following drawbacks: ultrasound examinations are easily influenced by subjective factors and are not sensitive enough to early lesions; most thyroid hormones and tumor markers such as AFP and CEA have poor sensitivity and specificity for early TC; tumor-specific nucleic acid characterization methods are relatively new and their clinical translation is still in its early stages, requiring more data to validate their effectiveness. Existing TC risk assessment systems suffer from high sample acquisition costs, difficulties, and low sample accumulation efficiency, making it challenging to iterate and upgrade existing TC risk assessment models. Summary of the Invention

[0006] To address the problems that existing TC risk assessment systems are not applicable to monitoring the risk of thyroid cancer in healthy individuals, and that high sample acquisition costs, difficulties, and low sample accumulation efficiency make it difficult to iterate and upgrade risk assessment models, this application provides a TC disease risk prediction modeling method based on cell proliferation markers.

[0007] This application provides a method for predicting and modeling the risk of TC disease based on cell proliferation markers, which adopts the following technical solution:

[0008] A modeling method for predicting the risk of thyroid cancer (TC) based on cell proliferation markers includes: collecting sample data; converting medical imaging examination results into quantitative TI-RADS grades, and setting TI-RADS grades of TR-3 and above as endpoint events; using multiple physical examination indicators, mainly serum thymidine kinase 1 (TK1) cell proliferation markers, as independent variables; and performing modeling through regression to obtain a model for predicting the risk of thyroid cancer (TC).

[0009] This application sets the intermediate state of TC (which is an intermediate state that must be passed through in the pathological process of TC and has a certain level of deterioration risk) as the endpoint event (that is, the set of imaging examination results with a certain risk of TC is set as the risk prediction endpoint event (TR-3, TR-4 and TR-5 are set as positive results)). Multiple physical examination indicators are used as independent variables. Since both independent and dependent variables are set at the physical examination stage, only physical examination results are required to develop the risk prediction model. There is no need to carry out a lot of inefficient follow-up work to collect clinical results, thereby reducing the cost and difficulty of obtaining positive result samples. The stable accumulation of positive samples becomes possible. Within the same time period, the feasibility of iterative upgrading of the risk prediction model obtained by this system is higher. Furthermore, by combining multiple health checkup indicators, primarily the results of serum thymidine kinase 1 (TK1), a cell proliferation marker, as independent variables, the resulting predictive model is better suited for predicting the risk of TC disease. Moreover, since this application uses multiple health checkup indicators as independent variables, the established TC disease risk prediction model can be widely applied to the general healthy population. For example, it can be used in health checkup institutions to predict the risk of TC disease in the general population, meeting the high-frequency and high-throughput requirements of universal screening.

[0010] Preferably, the modeling method using regression to obtain the risk prediction model for thyroid cancer (TC) specifically includes the following steps:

[0011] S1, using the three-fold cross-validation method to randomly sample and split the original sample data into a training set and a validation set;

[0012] S2, for different training sets, the imbalanced regression ILKL algorithm is repeatedly used to model and obtain multiple prediction models; the multiple prediction models obtained for the same training set form a prediction model library;

[0013] S3. Screen the prediction models in each prediction model library. If the prediction model formula contains the TK1 item independent variable, then include the prediction model in the initial screening prediction model group.

[0014] S4. For the prediction models in the initial screening prediction model group, use the corresponding validation set samples to validate the receiver operating characteristic curve (i.e., ROC curve) and calculate the area under the curve (AUC value).

[0015] S5. If the AUC value is greater than or equal to 0.7, then the prediction model in the corresponding initial screening prediction model group will be included in the final prediction model group.

[0016] S6. For each prediction model in the final prediction model group, optimize the model according to the parameter integration method to obtain the TC risk prediction model.

[0017] By adopting the above approach, especially by using the TK1 project independent variable and receiver operating characteristic curve to screen the prediction model, the TK1 test results can indicate the cell proliferation level and reflect the early abnormal proliferation changes of multiple types of tumor cells; the AUC value is greater than or equal to 0.7, proving that the prediction model has a certain degree of accuracy; in addition, by repeatedly using the unbalanced regression ILKL algorithm for modeling, the number of samples can be matched between highly imbalanced sample groups while maintaining the high representativeness of the samples, thus ultimately obtaining a more practical and accurate early risk prediction model for TC.

[0018] Preferably, the modeling using the imbalanced regression ILKL algorithm described in step S2 specifically includes the following steps:

[0019] S21, the samples are divided into two groups according to the endpoint event status. The samples with endpoint events of "general risk" (i.e., TR-1 and TR-2) are the "majority group", and the positive samples (i.e., TR-3 and above TI-RADS classification) are the minority group.

[0020] S22, set the physical examination indicators as independent variables as cluster indicators, set the endpoint event as sample labels, and set K (a natural number greater than 2 and less than or equal to 10) as the number of class groups;

[0021] S23, randomly select K samples from the multi-array sample as centroids, calculate the Euclidean distance between each point in the multi-array sample and each centroid sample based on the clustering index variable, and assign it to the set of the centroid with the shortest distance.

[0022] S24, after all data are grouped into K sets, the centroid of each set is recalculated;

[0023] S25. If the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold (i.e., the convergence state is reached), then the obtained K sets are used as K-Means clustering groups; otherwise, continue the iteration until the final K-Means clustering groups are obtained.

[0024] S26. Randomly sample K-Means cluster groups according to the ratio of minority groups to majority groups;

[0025] S27, merge all the sampled samples and the minority group samples, and obtain the prediction model through binary logistic regression.

[0026] By adopting the above approach and using the Imbalanced Regression (ILKL) algorithm for modeling, the sample imbalance problem between positive and healthy samples in real physical examination data can be effectively solved. Samples with "general risk" as the endpoint event are selected as the "majority group," and positive samples (i.e., TI-RADS classification of TR-3 and above) are selected as the minority group. At this point, by randomly sampling from K categories according to the ratio of different endpoint event states (positive group / majority group), a "general risk" sample set with an equal number of positive samples can be obtained. Performing binary logistic regression under the condition of sample equivalence can effectively avoid the problem of the prediction model biasing towards the majority group results due to sample imbalance. Furthermore, it expands the diversity of the prediction model by utilizing the true proportion of the positive population.

[0027] Preferably, the model optimization according to the parameter synthesis method described in step S6 specifically includes the following steps:

[0028] S61, using each prediction model in the final prediction model group, perform receiver operating characteristic (ROC) curve validation on all training and validation set data respectively and calculate the area under the curve (AUC) value.

[0029] S62, Sort the prediction models according to the prediction accuracy of positive samples in the verification results;

[0030] S63, Select the 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models:

[0031] The prediction model contains three parameters: project independent variables, coefficients of project independent variables, and constants. During optimization, all project independent variables in all models are combined to form the project independent variables of the final TC risk prediction model. The coefficients of each project independent variable in all models are summed and averaged to form the coefficients of the corresponding project independent variables in the TC risk prediction model. The constants in all models are summed and averaged to form the constants of the TC risk prediction model.

[0032] By adopting the above technical solutions, the accuracy of TC model predictions can be further improved.

[0033] Preferably, the following methods are used to screen other physical examination indicators that serve as independent variables:

[0034] First, (through significance analysis, correlation analysis, and rank sum analysis) verify the degree of strong correlation between the quantified transformation project and the endpoint event; if the significance of the correlation P < 0.1, then the project is included in the candidate projects; otherwise, it is not included in the candidate projects.

[0035] Secondly, for the candidate items obtained from the screening, the correlation between the items is checked (using the Box-Tidwell method, the correlation analysis between items, and the multicollinearity test). Items that meet the three assumptions of the independent variable are selected and included in the regression items. Finally, the results of the serum cytoplasmic thymidine kinase 1 test are used as independent variables. The three assumptions of the independent variable are that there is a linear relationship between the item and the logit transformation value of the endpoint event, and there is no multicollinearity and significant correlation between the items.

[0036] By employing the above technical solutions (especially significance analysis, correlation analysis, and rank-sum analysis), it can be ensured that there is a significant or relatively significant correlation between the detected items and the occurrence of the endpoint event (e.g., verified by the Box-Tidwell method), a linear relationship exists between the logit transformed values ​​of continuous independent variables and dependent variables, and (e.g., through inter-item correlation analysis and multicollinearity tests) the independence of the selected variable items is guaranteed. Including independent variables that meet these conditions in the "regression items" ensures that the differences in univariate analysis of independent variables are statistically significant, that there is a linear relationship between the logit transformed values ​​of continuous independent variables and dependent variables, and also avoids the distortion of the influence weight of certain factors due to the inclusion of highly correlated independent variable items.

[0037] Preferably, the medical imaging examination results are converted into quantitative TI-RADS grading by extracting the TI-RADS grading quantitative results from the text results of the medical imaging examination; after converting the medical imaging examination results into quantitative TI-RADS grading, the method further includes: assigning natural number values ​​to the classification option results, thereby converting the initial test and survey results data in different formats into consistent quantitative data.

[0038] By adopting the above technical solutions, medical imaging examination results and initial test and survey results data in different formats are converted into consistent quantitative data, which facilitates their use in modeling.

[0039] A further preferred method involves using an endpoint event extraction tool to extract TI-RADS grading and quantification results from textual results of medical imaging examinations, specifically including the following steps:

[0040] a. Using the LEN and SUBSTITUTE commands, output the number of times "TI" appears in the radiographic text results;

[0041] b. Use the FIND command to obtain the location value N1 of the first occurrence of "TI" in the text, and use the MID command to capture the hierarchical string after TI-RADS at position N1;

[0042] c. Using the first occurrence of the location value N1+1 as the starting position, continue using the FIND command to obtain the location value N2 of the second occurrence of "TI", and capture the TI-RADS hierarchical string at position N2;

[0043] d, take the second occurrence of the location value N2+1 as the starting position, continue to use the FIND command to locate, and so on, to obtain the hierarchical string after TI-RADS for all positions;

[0044] e. Use the VALUE and IFERROR commands to convert all strings obtained at all positions into numbers (or assign the value "0" if there are no numbers);

[0045] f, use the MAX command to output the maximum value L in the converted number;

[0046] g. Use the IF command to convert the obtained T value into a risk rating of 1 or 2.

[0047] By adopting the above technical solutions, the endpoint event extraction tool can be used to specifically discover TI-RADS classification information from massive amounts of medical image descriptive text, and the output can be converted according to the maximum value of its classification result, thereby improving conversion efficiency and reducing error rate.

[0048] Preferably, the process of assigning natural numbers to the classification options results involves using a logical value conversion tool to convert the items into quantitative values ​​according to whether they are binary, ordered, or unordered (for example, assigning a negative result of a binary item to 1 and a positive result to 2; converting multi-class items into a sequence of natural numbers (such as 1, 2, up to n) in order).

[0049] By adopting the above technical solutions, the projects are quantified according to different categories of binary classification, ordered and unordered, which is conducive to using the physical examination index data for modeling.

[0050] In the aforementioned TC disease risk prediction modeling method based on cell proliferation markers, the TC risk group information of the examinee is output by segmenting the risk probability results (e.g., segmenting according to CUTOFF=0.500, with those less than 0.5 being low-risk individuals and those greater than or equal to 0.5 being high-risk individuals).

[0051] By using the above methods, examinees can be divided into high-risk and low-risk groups, which helps to allocate most medical resources to the high-risk group.

[0052] In summary, this application includes at least one of the following beneficial technical effects:

[0053] 1. This application sets the intermediate state of TC (which is an intermediate state that must be passed through in the pathological process of TC and has a certain level of deterioration risk) as the endpoint event (that is, the comprehensive classification of imaging examination results with a certain risk of TC is set as the risk prediction endpoint event (TR-3, TR-4 and TR-5 are set as positive results)). Multiple physical examination indicators are used as independent variables. Since both independent and dependent variables are set at the physical examination stage, only physical examination results are required to develop the risk prediction model. There is no need to carry out a lot of inefficient follow-up work to collect clinical results, thereby reducing the cost and difficulty of obtaining positive result samples. The stable accumulation of positive samples becomes possible. Within the same time period, the feasibility of iterative upgrading of the risk prediction model obtained by this system is higher. Furthermore, by combining multiple physical examination indicators, primarily serum thymidine kinase 1 (TK1) test results, as independent variables, the resulting predictive model is better suited for predicting TC risk. Moreover, since this application uses multiple physical examination indicators as independent variables, the established TC risk prediction model can be widely applied to the general healthy population. For example, it can be used in physical examination institutions to predict the risk of TC in ordinary physical examination populations, meeting the high frequency and high throughput requirements of universal screening.

[0054] 2. This application uses the Imbalanced Regression (ILKL) algorithm for modeling, which effectively solves the sample imbalance problem between positive and healthy samples in real physical examination data. Samples with "general risk" as the endpoint event are selected as the "majority group," and positive samples (i.e., TI-RADS classification of TR-3 and above) are selected as the minority group. By randomly sampling from K categories according to the ratio of different endpoint event states (positive group / majority group), a "general risk" sample set with an equal number of positive samples can be obtained. Binary logistic regression is performed under the condition of sample equivalence, which effectively avoids the problem of the prediction model biasing towards the majority group results due to sample imbalance. Furthermore, it expands the diversity of the prediction model by utilizing the true proportion of the positive population. Attached Figure Description

[0055] Figure 1 This is a flowchart of a method according to one embodiment of this application.

[0056] Figure 2 This is a flowchart illustrating the specific modeling method for the TC risk prediction model in this application.

[0057] Figure 3 This is a flowchart illustrating the project selection and ILKL algorithm modeling method used in this application.

[0058] Figure 4 A flowchart illustrating a method for extracting TI-RADS grading and quantification results from textual results of medical imaging examinations using an endpoint event extraction tool.

[0059] Figure 5 This is a schematic diagram of a block diagram for assessing the risk of TC using the model established in this application.

[0060] Figure 6 This is a schematic diagram of the ROC curve used to validate the established prediction model in the experimental example. Detailed Implementation

[0061] The following is in conjunction with the appendix Figure 1-6 This application will be described in further detail.

[0062] In existing technologies, such as the risk prediction model for Hashimoto's thyroiditis combined with differentiated thyroid cancer proposed by Wang Xiaoliang et al. in 2019, which assesses the TC risk of patients with Hashimoto's thyroiditis and differentiated thyroid cancer, this TC risk assessment system uses the final diagnosis of thyroid cancer as the endpoint event. Therefore, it requires thyroid cancer patients as the target sample collection. Since confirmed thyroid cancer patients actually account for only 0.021% of the screened population, sample acquisition is costly, difficult, and has low sample accumulation efficiency, making it very difficult to iterate and upgrade existing TC risk assessment models. As the core of the risk assessment system, the iterative upgrade of the clinical prediction model is a key way to improve the effectiveness of the risk assessment system. However, because physical examination data and clinical data are generally isolated into "information silos," the completeness of physical examination data, the consistency of clinical results, and the correlation between physical examination results and clinical results (large time intervals leading to a lack of correspondence) all interfere with the effectiveness of positive samples. As a result, the number of usable positive samples is far lower than the actual occurrence. This increases the difficulty of predictive model research, makes it difficult to obtain TC risk prediction models with application value, and also reduces the feasibility of iterating risk prediction algorithms. Based on this problem, the inventors conducted research and found that, according to the latest version of the ACR TI-RADS v2017 white paper, thyroid nodules detected by imaging are divided into five categories: benign (TR-1), highly likely benign (TR-2), suspected TC (TR-3), moderately suspected TC (TR-4), and highly suspected TC (TR-5). According to a 2019 report by Wildman-Tobriner B et al., the positive predictive value (PPV) for TC is 0% for TR-1; 1.5% for TR-2; and 11.8% for TR-3. The PPV values ​​for TC in TR-4 and TR-5 are 29.4% and 57.4%, respectively. Therefore, TR-1 and TR-2 results can be considered as a general risk status for TC, while TR-3, TR-4, and TR-5 results can be considered as a high-risk status for TC.

[0063] Therefore, the inventors creatively conceived of setting the intermediate state of TC (which is an intermediate state that must be passed through in the pathological process of TC and has a certain level of deterioration risk characteristics) which can be obtained from physical examination samples as the endpoint event (that is, converting medical imaging examination results into quantitative endpoint event states, setting TR-3 and above TI-RADS grades as endpoint events (that is, setting TR-3, TR-4 and TR-5 as positive results)). Multiple physical examination indicators are used as independent variables. Since both independent and dependent variables are set at the physical examination stage, only physical examination results are required to develop a risk prediction model. There is no need to carry out a lot of inefficient follow-up work to collect clinical results, which reduces the cost and difficulty of obtaining positive result samples. Stable accumulation of positive samples becomes possible. Within the same time period, the system has a higher feasibility for iterative upgrades of the risk prediction model. Furthermore, by combining multiple health checkup indicators, primarily serum thymidine kinase 1 (TK1) test results, as independent variables, the resulting predictive model is better suited for predicting TC risk. Moreover, the use of multiple health checkup indicators as independent variables in this application allows the established TC risk prediction model to be widely applicable to the general healthy population. For example, it can be used in health checkup institutions to predict TC risk among ordinary individuals undergoing routine health checkups, meeting the high-frequency and high-throughput requirements of widespread screening. In other words, by changing the data acquisition model and reducing the cost of obtaining positive samples, this application effectively improves the development efficiency of the TC risk prediction model. Based on a long-term, stable accumulation of positive samples, the TC risk prediction model can undergo more thorough algorithm iterations, resulting in a more stable and efficient TC risk assessment system, ultimately achieving highly efficient allocation of medical resources to high-risk TC populations.

[0064] This application discloses a method for predicting and modeling the risk of TC disease based on cell proliferation markers. (Refer to...) Figure 1 A method for predicting the risk of TC disease based on cell proliferation markers, comprising:

[0065] Step 1: Collect sample data;

[0066] Step 2: Convert the medical imaging results into quantitative TI-RADS classifications and set TR-3 and above TI-RADS classifications as endpoint events; use various physical examination indicators, mainly serum thymidine kinase 1 (TK1) test results, as independent variables.

[0067] Step 3: Model the disease using regression methods (such as the general formula of binary logistic regression risk model) to obtain a thyroid cancer (TC) risk prediction model.

[0068] In practice, the process begins with an overall observation of the physical examination results. Results from the same institution and testing system are selected. Then, endpoint event extraction tools and logistic value conversion tools are used to convert the different forms of the original items into a unified quantitative form. Next, a data cleaning process removes samples with data integrity and standardization deficiencies (e.g., samples with no TK1 value or TI-RADS results, and samples whose converted quantitative results significantly exceed the measurement range, such as negative numbers or outlier samples (outlier determination follows the method in 6.2.1 "Upper Case" of GB / T 4883-2008 Statistical Processing and Interpretation of Data—Judgment and Processing of Outliers in Normal Samples)). Finally, an unstructured database can be used to store the original state and numerically converted results, and the data can be output as analyzable data files according to analytical requirements.

[0069] Generally, physical examination results are divided into three categories according to the form of the results: 1) medical imaging examination results, including ultrasound / contrast ultrasound (CEUS) items; 2) continuous value items; 3) classification result items, including binary items, ordered multi-class items, and unordered multi-class items.

[0070] Optionally, the modeling method described above uses regression analysis to obtain a predictive model for the risk of thyroid cancer (TC), such as... Figure 2 As shown, the specific steps include:

[0071] S1, using the three-fold cross-validation method to randomly sample and split the original sample data into a training set and a validation set;

[0072] S2, for different training sets, the imbalanced regression ILKL algorithm is repeatedly used to model and obtain multiple prediction models; the multiple prediction models obtained for the same training set form a prediction model library;

[0073] S3. Screen the prediction models in each prediction model library. If the prediction model formula contains the TK1 item independent variable, then include the prediction model in the initial screening prediction model group.

[0074] S4. For the prediction models in the initial screening prediction model group, use the corresponding validation set samples to validate the receiver operating characteristic curve (i.e., ROC curve) and calculate the area under the curve (AUC value).

[0075] S5. If the AUC value is greater than or equal to 0.7, then the prediction model in the corresponding initial screening prediction model group will be included in the final prediction model group.

[0076] S6. For each prediction model in the final prediction model group, optimize the model according to the parameter integration method to obtain the TC risk prediction model.

[0077] In practice, for example, if the sample size for model development is 5259, it is divided into three groups using a three-fold sampling method: Group 1 (1753 samples), Group 2 (1753 samples), and Group 3 (1753 samples). The ILKL algorithm is used for sample matching during model development, and the ROC curve validation method is used for model selection. In the three-fold combination, Group 1 is used as the validation set, and Groups 2 and 3 are merged to form the training set. Figure 3 By filtering the items in the training set and iteratively processing them using the ILKL algorithm, we obtained prediction model library 1 (containing 100 prediction models). Then, we repeated the above steps in combination two and combination three, which were sampled using the three-fold method, to form different prediction model libraries.

[0078] Optional, such as Figure 3 As shown, the modeling using the imbalanced regression ILKL algorithm described in step S2 specifically includes the following steps:

[0079] S21, the samples are divided into two groups according to the endpoint event status. The samples with endpoint events of "general risk" (i.e., TR-1 and TR-2) are the "majority group", and the positive samples (i.e., TR-3 and above TI-RADS classification) are the minority group.

[0080] S22, set the physical examination indicators as independent variables as cluster indicators, set the endpoint event as sample labels, and set K (a natural number greater than 2 and less than or equal to 10) as the number of class groups;

[0081] S23, randomly select K samples from the multi-array sample as centroids, calculate the Euclidean distance between each point in the multi-array sample and each centroid sample based on the clustering index variable, and assign it to the set of the centroid with the shortest distance.

[0082] S24, after all data are grouped into K sets, the centroid of each set is recalculated;

[0083] S25. If the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold (i.e., the convergence state is reached), then the obtained K sets are used as K-Means clustering groups; otherwise, continue the iteration until the final K-Means clustering groups are obtained.

[0084] S26. Randomly sample K-Means cluster groups according to the ratio of minority groups to majority groups;

[0085] S27, merge all the sampled samples and the minority group samples, and obtain the prediction model through binary logistic regression.

[0086] If there are 120 positive samples, in the sampling of the healthy group, random sampling is carried out according to the ratio of positive group to healthy group in K-Means clustering, so that the number of samples after sampling the healthy group is close to 120, thereby achieving the purpose of reducing prediction bias.

[0087] Optionally, the model optimization according to the parameter synthesis method described in step S6 specifically includes the following steps:

[0088] S61, using each prediction model in the final prediction model group, perform receiver operating characteristic (ROC) curve validation on all training and validation set data respectively and calculate the area under the curve (AUC) value.

[0089] S62, Sort the prediction models according to the prediction accuracy of positive samples in the verification results;

[0090] S63, Select the 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models:

[0091] The prediction model contains three parameters: project independent variables, coefficients of project independent variables, and constants. During optimization, all project independent variables in all models are combined to form the project independent variables of the final TC risk prediction model. The coefficients of each project independent variable in all models are summed and averaged to form the coefficients of the corresponding project independent variables in the TC risk prediction model. The constants in all models are summed and averaged to form the constants of the TC risk prediction model.

[0092] For example, when performing parameter optimization, the prediction model contains three parameters, A n These are fixed numbers for all independent variable items in the model, actually numbered from A1 to A2. n B m-n The project coefficient refers to A in model M (e.g., M = 1 to 5). n The project parameter, actually numbered B 1-1 To B 5-n If any model does not contain a certain A n The project's independent variable, with its parameter B value set to 0; C M These are constants in model M, actually numbered C1-C5. The optimized model contains all A values ​​from the five merged models. n The optimization of parameters B and C follows the formula below:

[0093] C op = (C1+C2+C3+C4+C5) / 5

[0094] B op(m-n) = (B 1-n +B 2-n +B3-n +B 4-n +B 5-n ) / 5.

[0095] Optional, such as Figure 3 As shown, other physical examination indicators used as independent variables were screened using the following methods: First, the degree of close correlation between the quantified transformed items and the endpoint event was verified (through significance analysis, correlation analysis, and rank sum analysis); if the significance of the correlation P < 0.1, the item was included in the candidate items; otherwise, it was not included in the candidate items.

[0096] Secondly, for the candidate items obtained from the screening, the correlation between the items is checked (using the Box-Tidwell method, inter-item correlation analysis, and multicollinearity test). Items that meet the three assumptions of the independent variable are selected and included in the regression project. Finally, the results of serum cytoplasmic thymidine kinase 1 detection are used as independent variables. The three assumptions of the independent variable are that there is a linear relationship between the item and the transformed value of the endpoint event (logit), and there is no multicollinearity and significant correlation between the items.

[0097] During sample collection, the following information is collected from the examinee: 1) Basic information of the examinee (sensitive information removed); 2) Family medical history / personal medical history; 3) Lifestyle survey results; 4) Biochemical test results; 5) Medical imaging examination results. The physical examination items and the converted quantitative items are then imported into the resulting database storage system. According to the logical relationship of the recorded content, the database is divided into basic information of the examinee, family medical history / personal medical history, lifestyle survey results, individual test results, and medical imaging examination results. The individual test results are further composed of blood tests, biochemical tests, infection and immune tests, and body fluid tests.

[0098] Optionally, the medical imaging examination results can be converted into quantitative TI-RADS grading by extracting the TI-RADS grading quantitative results from the text results of the medical imaging examination. After converting the medical imaging examination results into quantitative TI-RADS grading, the method further includes: assigning natural number values ​​to the classification option results, thereby converting the initial test and survey results data in different formats into consistent quantitative data.

[0099] Optionally, the TI-RADS grading and quantification results from the textual results of medical imaging examinations can be extracted using endpoint event extraction tools, such as... Figure 4 As shown, the specific steps include:

[0100] a. Using the LEN and SUBSTITUTE commands, output the number of times "TI" appears in the radiographic text results;

[0101] b. Use the FIND command to obtain the location value N1 of the first occurrence of "TI" in the text, and use the MID command to capture the hierarchical string after TI-RADS at position N1;

[0102] c. Using the first occurrence of the location value N1+1 as the starting position, continue using the FIND command to obtain the location value N2 of the second occurrence of "TI", and capture the TI-RADS hierarchical string at position N2;

[0103] d, take the second occurrence of the location value N2+1 as the starting position, continue to use the FIND command to locate, and so on, to obtain the hierarchical string after TI-RADS for all positions;

[0104] e. Use the VALUE and IFERROR commands to convert all strings obtained at all positions into numbers (or assign the value "0" if there are no numbers);

[0105] f, use the MAX command to output the maximum value L in the converted number;

[0106] g. Use the IF command to convert the obtained T value into a risk rating of 1 or 2.

[0107] In practice, for example, the command line in Excel 2013 can be used to extract the TI-RADS grading and quantification results from the text results of medical imaging examinations:

[0108] Step 1):

[0109] fx = 0.5 * (LEN(initial text) - LEN(SUBSTITUTE(initial text,"TI","")));

[0110] Step 2):

[0111] fx = MID(A3, FIND("TI", initial text) + 8, 1);

[0112] Step 3):

[0113] fx = MID(A3, FIND("TI", initial text, N1+1) + 8, 1);

[0114] Step 4):

[0115] fx = MID(A3, FIND("TI", initial text, N2+1) + 8, 1);

[0116] Step 5):

[0117] fx = MID(A3, FIND("TI", text table, N3+1) + 8, 1);

[0118] Step 6):

[0119] fx = IFERROR(VALUE("fetch TI-RADS hierarchical string"), 0);

[0120] Step 7):

[0121] fx = MAX(Step 2 extracts grade values: Step 5 extracts grade values);

[0122] Step 8):

[0123] fx = IF(L>2,2,1).

[0124] Optionally, the process of assigning natural numbers to the classification options results involves using a logical value conversion tool to convert the items into quantitative values ​​according to whether they are binary, ordered, or unordered. For example, binary classification items are converted to 0 / 1 or 1 / 2 according to positive / negative results, such as male / female being converted to 1 / 2 respectively; TR-1 and TR-2 are converted to 1, and TR-3, TR-4, and TR-5 are uniformly converted to 2; multi-classification items are converted into a natural number sequence (such as 1, 2, up to n) in order.

[0125] In this application, the TC risk group information of the examinee can be output by segmenting the risk probability results (e.g., segmenting according to CUTOFF=0.500, with those less than 0.5 being low-risk groups and those greater than or equal to 0.5 being high-risk groups).

[0126] In practice, long-term health checkup data from the same location and institution can be collected, including TK1 testing and other routine health checkup items (including but not limited to blood tests, tumor markers, etc.). Figure 5 As shown, according to the TC risk prediction model established in this application, specific items in the physical examination report are selected; the test results of the TK1 reagent kit are imported into the prediction model; the risk probability is calculated and judged through the model; and the prediction results are output in text format based on the calculation results. The TC disease risk prediction model established in this application is applicable to ordinary healthy people and can be used in physical examination institutions for TC disease risk prediction.

[0127] Experimental example:

[0128] Using physical examination data from a hospital from 2009 to 2014 (total sample size 20,679), including 16,480 individuals in the general risk group (TR-1 and TR-2) and 4,199 individuals in the negative group (reflecting the true proportions and improving the scalability of the regression model and the practical significance of the risk assessment method), the model was developed based on the modeling method proposed in this application. The final predictive model used age stratification, gender, and eight biomarkers or physical examination items as independent variables: serum thymidine kinase 1 concentration (TK1), α-hydroxybutyrate dehydrogenase, high-density lipoprotein, uric acid, alanine aminotransferase (ALT), platelet count, carcinoembryonic antigen (CEA), and aspartate aminotransferase (AST).

[0129] Age is expressed in years; males are assigned a value of 1 and females a value of 0; TK1 value (pM) is measured using a CIS series chemiluminescence digital imaging analyzer and a thymidine kinase 1 diagnostic kit manufactured by Shenzhen Huarui Tongkang Biotechnology Co., Ltd.; uric acid, platelet count, and high-density lipoprotein can be measured using an XFA6100 fully automated blood cell analyzer; α-hydroxybutyrate dehydrogenase, aspartate aminotransferase, and alanine aminotransferase can be measured using an MD-100 semi-automatic biochemical analyzer; carcinoembryonic antigen (CEA) can be measured using an ELISA kit (supplier: Shanghai Guandao Biotechnology Co., Ltd.) and a fully automated biochemical analyzer.

[0130] The corresponding function formula for the risk prediction model of thyroid cancer (TC) is: Risk probability P = ExpX / (1 + ExpX), where X = -1.972 + (0.0285 × age (years)) + (-0.472 × gender assignment) + (0.23 × TK1 (pM)) + (0.004 × platelet count (10)). 9 / L))+(-1.0485×high-density lipoprotein(mmol / L))+(-0.002×uric acid(μmol / L))+(0.021×α-hydroxybutyrate dehydrogenase(U / L))+(-0.003×alanine aminotransferase(U / L))+(-0.011×aspartate aminotransferase(U / L))+(0.1745×CEA(μg / L)).

[0131] The risk probability P represents the likelihood of a TI-RADS 3 outcome in medical imaging. That is, the TI-RADS 3 classification result in the thyroid imaging reporting and data system is used as the endpoint event of the prediction model, anchoring a definite TC risk probability. When the risk probability P < 0.5, it can be judged as a general risk; while when the risk probability P ≥ 0.5, it can be judged as a high risk.

[0132] In addition, the predictive model was validated using physical examination data from a hospital from 2009 to 2017 (total sample size 4552). Figure 6As shown, the AUC value was 0.749, with a 95% CI of 0.736–0.761. The Youden index obtained by using the bootstrap sampling confidence interval (1000 iterations; random number seed: 978) test was 0.3873, with a sensitivity of 55.53 and a specificity of 84.20 (threshold 0.5).

[0133] In addition, the inventors also compared the value of the prediction model with the combined detection method. The results showed that the TC risk prediction model established in this experimental case, which uses 10 items for simultaneous detection, has a higher prediction accuracy than the cell proliferation marker TK1 or the tumor marker CEA used alone or in combination with TK1 and CEA.

[0134] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the methods and principles of this application should be covered within the scope of protection of this application.

Claims

1. A method for predicting and modeling the risk of TC disease based on cell proliferation markers, characterized in that, include: Collect sample data; Medical imaging results were converted into quantitative TI-RADS grading, and TI-RADS grading of TR-3 and above was set as the endpoint event; multiple physical examination indicators, mainly the cell proliferation marker serum thymidine kinase 1 (TK1) test results, were used as independent variables. A regression model was used to model the risk of thyroid cancer (TC). The aforementioned modeling using regression to obtain a risk prediction model for thyroid cancer (TC) includes the following steps: S1, using the three-fold cross-validation method to randomly sample and split the original sample data into a training set and a validation set; S2, for different training sets, multiple prediction models are obtained by repeatedly using the imbalanced regression ILKL algorithm; multiple prediction models obtained for the same training set form a prediction model library; S3. Screen the prediction models in each prediction model library. If the prediction model formula contains the TK1 item independent variable, then include the prediction model in the initial screening prediction model group. S4. For the prediction models in the initial screening prediction model group, use the corresponding validation set samples to perform receiver operating characteristic curve validation and calculate the area under the curve (AUC). S5. If the AUC value is greater than or equal to 0.7, then the prediction model in the corresponding initial screening prediction model group will be included in the final prediction model group. S6. For each prediction model in the final prediction model group, optimize the model according to the parameter integration method to obtain the TC risk prediction model. Step S6, which describes model optimization using a parameter synthesis approach, specifically includes the following steps: S61, using each prediction model in the final prediction model group, perform receiver operating characteristic curve validation on all training and validation set data and calculate the area under the curve (AUC) value. S62, Sort the prediction models according to the prediction accuracy of positive samples in the verification results; S63, Select the 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models: The prediction model contains three parameters: project independent variables, coefficients of project independent variables, and constants. During optimization, all project independent variables in all models are combined to form the project independent variables of the final TC risk prediction model. The coefficients of each project independent variable in all models are summed and averaged to form the coefficients of the corresponding project independent variables in the TC risk prediction model. The constants in all models are summed and averaged to form the constants of the TC risk prediction model.

2. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 1, characterized in that, Step S2, which describes modeling using the imbalanced regression ILKL algorithm, specifically includes the following steps: S21, the samples are divided into two groups according to the state of the endpoint event. The sample group with a general risk endpoint event is the majority group, and the positive sample group is the minority group. S22, set the physical examination indicators as independent variables as clustering indicators, set the endpoint event as the sample label, and set K as the number of class groups; S23, randomly select K samples from the multi-array sample as centroids, calculate the Euclidean distance between each point in the multi-array sample and each centroid sample based on the clustering index variable, and assign it to the set of the centroid with the shortest distance. S24, after all data are grouped into K sets, the centroid of each set is recalculated; S25. If the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold, then the obtained K sets are used as K-Means clustering groups; otherwise, continue the iteration until the final K-Means clustering groups are obtained. S26. Randomly sample K-Means cluster groups according to the ratio of minority groups to majority groups; S27, merge all the sampled samples and the minority group samples, and obtain the prediction model through binary logistic regression.

3. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 1, characterized in that, The following methods were used to screen other physical examination indicators that were used as independent variables: First, verify the degree of strong correlation between the quantified conversion project and the endpoint event; if the significance of the correlation P<0.1, the project is included in the candidate projects; otherwise, it is not included in the candidate projects. Secondly, for the candidate items obtained from the screening, the correlation between the items is investigated, and the items that meet the three assumptions of the independent variable are selected and included in the regression items. Finally, the results of serum thymidine kinase 1 detection are used as independent variables. The three assumptions of the independent variable are that there is a linear relationship between the item and the logit transformation value of the endpoint event, and there is no multicollinearity and significant correlation between the items.

4. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 1, characterized in that, By extracting the TI-RADS grading quantification results from the textual results of medical imaging examinations, the medical imaging examination results can be converted into quantified TI-RADS grading. After converting medical imaging examination results into quantitative TI-RADS classifications, the process also includes: assigning natural numbers to the classification option results, thereby converting the initial test and survey results data in different formats into consistent quantitative data.

5. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 4, characterized in that, The endpoint event extraction tool was used to extract TI-RADS grading and quantification results from textual results of medical imaging examinations. Includes the following steps: a. Using the LEN and SUBSTITUTE commands, output the number of times "TI" appears in the image text results; b. Use the FIND command to obtain the location value N1 of the first occurrence of "TI" in the text, and use the MID command to capture the hierarchical string after TI-RADS at position N1; c. Using the first occurrence of the location value N1+1 as the starting position, continue using the FIND command to obtain the location value N2 of the second occurrence of "TI", and capture the TI-RADS hierarchical string at position N2; d, take the second occurrence of the location value N2+1 as the starting position, continue to use the FIND command to locate, and so on, to obtain the hierarchical string after TI-RADS for all positions; e. Use the VALUE and IFERROR commands to convert all strings obtained at all positions into numbers; f, use the MAX command to output the maximum value T in the converted number; g. Use the IF command to convert the obtained T value into a risk level of 1 or 2.

6. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 4, characterized in that, The process of assigning natural numbers to the classification options results involves using a logical value conversion tool to convert the items into quantitative values ​​according to their binary, ordered, and unordered categories.

7. The TC disease risk prediction modeling method based on cell proliferation markers according to claim 1, characterized in that, By identifying segments of the risk probability results, the TC risk cluster information of the examinee is output.

Citation Information

Patent Citations

  • Early diagnosis method and diagnostic kit for liver cancer

    CN109557312A

  • Multi-physiological data fusion analysis method based on neural network

    CN110378394A

  • Disease suffering risk prediction method and device

    CN110838366A