A disease risk prediction modeling method based on LI-RADS grading

By setting the HCC intermediate state as the endpoint event in the physical examination sample and combining the TK1 detection results, an HCC disease risk prediction model based on LI-RADS classification was established, which solved the problem that the HCC risk assessment system in the prior art was difficult to apply to healthy groups, and efficient sample accumulation and model iterative upgrade was achieved.

CN114550921BActive Publication Date: 2025-05-06SHENZHEN HUARUI TONGKANG BIOTECHNOLOGICAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202011343525.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-26
Publication Date
2025-05-06
Estimated Expiration
2040-11-26

AI Technical Summary

Technical Problem

The existing HCC risk assessment system cannot be applied to healthy groups for risk monitoring of hepatocellular carcinoma, and the sample acquisition cost is high, difficult, and low sample accumulation efficiency, making it difficult to iteratively upgrade the risk assessment model.

Method used

A disease risk prediction modeling method based on LI-RADS grading was adopted. By setting the HCC intermediate state in the physical examination sample as the endpoint event, and combining the serum cytoplasmic thymidine kinase 1 (TK1) detection results as the main independent variable, a regression method was used to establish a risk prediction model of HCC disease by hepatocellular carcinoma.

Benefits of technology

It reduces the cost and difficulty of obtaining positive results samples, stabilizes the accumulation of positive samples, improves the feasibility of iterative upgrade of the risk prediction model, and the model can be widely used in the prediction of HCC disease risk in healthy people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 201126095136
    Figure 201126095136
  • Figure 201126095147
    Figure 201126095147
  • Figure 201126095159
    Figure 201126095159
Patent Text Reader

Abstract

The present application relates to a disease risk prediction modeling method based on LI-RADS grading, including: collecting sample data; converting medical imaging examination results into quantitative endpoint event states, and setting LR-3 and above LI-RADS grades as endpoint events; using multiple physical examination index items mainly based on serum cytoplasmic thymidine kinase 1, i.e., TK1 test results as independent variables; modeling through regression method, thereby obtaining a hepatocellular carcinoma HCC disease risk prediction model. The present application sets a set of imaging examination results with a determined HCC risk as a risk prediction endpoint event, thereby reducing the cost and difficulty of obtaining positive result samples, making it possible to accumulate positive samples stably, and in the same time period, the iterative upgrade feasibility of the risk prediction model obtained by the system is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of disease risk prediction modeling, and in particular to a disease risk prediction modeling method based on LI-RADS grading. Background Art

[0002] Liver cancer is one of the most common malignant tumors. According to statistics in 2018, liver cancer patients accounted for 4.7% of the total cancer population worldwide, with a mortality rate as high as 8.2%, especially in male cancer patients (mortality rate 10.2%), and lung cancer ranked first and second in mortality rate. Hepatocellular carcinoma (HCC) is the main type of primary liver cancer, accounting for 85%-90% of the total incidence. HCC has complex causes, no typical symptoms in the early stage, rapid pathological progression, and significant differences in prognosis at different stages. The 5-year survival rate for early diagnosis is 70%, and the 5-year survival rate for late diagnosis is less than 5%. The number of people with primary liver cancer in China accounts for about 50% of the world, and the number is increasing year by year. The 5-year survival rate is only 14.1%. One of the main reasons is that 70-85% of patients are already in the middle and late stages when diagnosed, and have lost the opportunity for surgery. Therefore, strengthening prevention and universal screening of high-risk groups for HCC is of great practical significance.

[0003] In recent years, the technologies for screening high-risk populations for HCC mainly include: 1) ultrasound combined with serum alpha-fetoprotein (AFP); 2) other tumor markers, such as alpha-fetoprotein isoform (AFP-L3), phosphatidylinositol glycan 3 (GPC3), osteopontin (OPN), des-γ-carboxyprothrombin (DCP), Golgi protein 73 (GP73), glycoprotein DKK1, α-L-fucosidase (AFU), carbohydrate antigen 19-9 (CA19-9), etc.; 3) tumor characteristic nucleic acid identification methods, such as circulating free microRNA (cfRNA), circulating tumor DNA (ctDNA), DNA methylation identification, second-generation sequencing, etc.; 4) tumor cell identification methods, such as circulating tumor cells (CTC), etc.; 5) imaging examination methods, such as X-ray computed tomography (CT), magnetic resonance imaging (MRI), ultrasound angiography (CEUS), hepatocyte-specific contrast agent enhanced scanning (EOB-MRI), etc.

[0004] Developing clinical prediction models and building risk assessment systems to test the risk of serious diseases is a hot topic in international epidemiology and oncology research. The main HCC-related risk assessment systems currently include: 1) Hepatitis B patient liver cancer risk prediction (CAMD), which uses age, gender, liver cirrhosis, and diabetes information to assess the risk of liver cancer in hepatitis B patients after 3 years; 2) Early HCC recurrence risk model after liver cancer resection, which assesses the risk of HCC recurrence through gender, albumin-bilirubin index, AFP, liver cancer size, and tumor number; 3) Chronic liver disease patient liver cancer risk prediction model (aMAP score), which assesses the 30-day mortality risk and 5-year cumulative liver cancer incidence of the subjects through age, gender, total bilirubin, albumin level, platelets, and blood creatinine level factors. Regarding the above-mentioned related technologies, the inventors believe that the following defects exist: ultrasound examination is easily affected by subjective factors and is not sensitive enough to early lesions; most of the biomarkers such as AFP, AFP-L3, OPN, etc. have poor sensitivity and specificity for early HCC; the tumor cell and tumor characteristic nucleic acid identification methods appeared relatively late and the clinical transformation progress is still early, so more data needs to be accumulated to verify the effectiveness of the methods; medical imaging methods have strict restrictions on the interval period, and multiple examinations in a short period of time are likely to increase the risk of genetic mutations in the subjects, and the time interval between two tests is too long. Due to the rapid progression of HCC itself, the problem of "interphase tumors" is prone to occur, and the instrument investment cost for universal screening of this method is also relatively high. A common limitation of existing HCC-related risk assessment systems is that the subjects are actually already in the clinical stage of liver disease, and the various indicators involved in the system are mostly test results in the clinical stage. This implicit basic condition does not meet the initial conditions of the healthy population that needs to undergo HCC risk assessment. In addition, the existing HCC risk assessment system has high sample acquisition costs, great difficulty, and low sample accumulation efficiency, which makes it very difficult to iterate and upgrade the existing HCC risk assessment model, which in turn leads to limited accuracy in risk prediction. Summary of the invention

[0005] In order to solve the problem that the existing HCC risk assessment system cannot be applied to healthy people for hepatocellular carcinoma risk monitoring, and that sample acquisition is costly and difficult, and sample accumulation efficiency is low, resulting in difficulty in iterative upgrading of risk assessment models, the present application provides a disease risk prediction modeling method based on LI-RADS grading.

[0006] A disease risk prediction modeling method based on LI-RADS grading includes: collecting sample data; converting a set of medical imaging examination results into a quantitative endpoint event state, and setting LI-RADS grading of LR-3 and above as the endpoint event; using multiple physical examination index items mainly including serum cytoplasmic thymidine kinase 1 (TK1) test results as independent variables; and modeling through regression method to obtain a hepatocellular carcinoma (HCC) disease risk prediction model.

[0007] The present application sets the intermediate state of HCC that can be obtained in the physical examination sample (the intermediate state of HCC is an intermediate state that must be passed in the pathological process of HCC and has a certain level of deterioration risk characteristics) as the endpoint event (that is, the set of imaging examination results with a certain HCC risk is set as the risk prediction endpoint event (LR-3, LR-4 and LR-5 are set as positive results)), and multiple physical examination indicator items as independent variables. Since both the independent variables and the dependent variables are set in the physical examination stage, the risk prediction model can be developed by providing only the physical examination results, and there is no need to conduct a large number of inefficient follow-up work to collect clinical results, thereby reducing the cost and difficulty of obtaining positive result samples, and the stable accumulation of positive samples becomes possible. In the same time period, the iterative upgrade feasibility of the risk prediction model obtained by the system is higher. In addition, by combining multiple physical examination indicator items, mainly serum cytoplasmic thymidine kinase 1, i.e., TK1 test results, as independent variables, the prediction model obtained by modeling can be better applicable to the prediction of HCC risk; and in this application, multiple physical examination indicator items are used as independent variables, so the established HCC risk prediction model can be generally applicable to the general healthy population, such as being used in physical examination institutions to predict the HCC risk of ordinary physical examination populations, and can adapt to the high frequency and high throughput requirements of general screening.

[0008] Preferably, the modeling by regression method to obtain a risk prediction model for hepatocellular carcinoma (HCC) specifically includes the following steps:

[0009] S1, use three-fold cross-validation method to randomly sample and split the original sample data into training set and validation set;

[0010] S2, for different training sets, multiple prediction models are obtained by repeatedly using the unbalanced regression ILKL algorithm; multiple prediction models obtained corresponding to the same training set form a prediction model library;

[0011] S3, screening the prediction models in each prediction model library, if the prediction model formula contains the TK1 project independent variable, then the prediction model is included in the initial screening prediction model group;

[0012] S4, for the prediction model in the initial screening prediction model group, the corresponding validation set samples were used to perform receiver operating characteristic curve (ROC curve) validation and calculate the area under the curve (AUC) value;

[0013] S5, if the AUC value is greater than or equal to 0.7, the prediction model in the corresponding preliminary screening prediction model group is included in the final prediction model group;

[0014] S6. Optimize each prediction model in the final prediction model group according to a parameter comprehensive method to obtain an HCC risk prediction model.

[0015] By adopting the above scheme, especially by using the TK1 project independent variable and the receiver operating characteristic curve to screen the prediction model, the TK1 test results can indicate the level of cell proliferation and reflect the early abnormal proliferation changes of multiple types of tumor cells; the ACU value is greater than or equal to 0.7, proving that the prediction model has a certain accuracy; in addition, by repeatedly using the unbalanced regression ILKL algorithm for modeling, it is possible to match the number between highly unbalanced sample groups while maintaining the high representativeness of its samples, thereby ultimately obtaining a more practical and accurate HCC early risk prediction model.

[0016] Preferably, the modeling using the unbalanced regression ILKL algorithm described in step S2 specifically includes the following steps:

[0017] S21, samples were divided into two groups according to the endpoint event status, where samples with “average risk” (i.e., LR-1 and LR-2) as the “majority group” and positive samples (i.e., LI-RADS grade LR-3 and above) as the minority group;

[0018] S22, setting the physical examination index item as the independent variable as the clustering index variable, setting the endpoint event as the sample label, and setting K (a natural number greater than 2 and less than or equal to 10) as the number of category groups;

[0019] S23, randomly selecting K samples from the multi-array samples as centroids, for each point in the multi-array samples, calculating the Euclidean distance between it and each centroid sample based on the clustering indicator variable, and dividing it into the set to which the centroid with the shortest distance belongs;

[0020] S24, after all data are classified into K sets, the centroid of each set is recalculated;

[0021] S25, if the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold (i.e., convergence is reached), the obtained K sets are used as K-Means clustering groups; otherwise, the iteration is continued until the K-Means clustering groups are finally obtained;

[0022] S26, random sampling of K-Means clustering groups according to the ratio of minority group to majority group;

[0023] S27, merge all the sampled samples with the minority group samples, and obtain the prediction model through binary logistic regression.

[0024] By adopting the above scheme and using the unbalanced regression ILKL algorithm for modeling, the sample imbalance problem between positive samples and healthy samples in real physical examination data can be effectively solved. The samples with the endpoint event of "general risk" are selected as the "majority group", and the positive sample group (i.e., LI-RADS grade of LR-3 and above) is selected as the minority group. At this time, as long as the ratio of different states of the endpoint event (positive group / majority group) is randomly sampled in K categories, a "general risk" sample set equal to the number of samples in the positive group can be obtained; binary logistic regression is performed under the condition of sample equality, which can effectively avoid the problem that the prediction model is biased towards the majority group result due to sample imbalance, and on the other hand, the diversity of the prediction model is expanded through the actual proportion of the positive population.

[0025] Preferably, the model optimization according to the parameter comprehensive method described in step S6 specifically includes the following steps:

[0026] S61, using each prediction model in the final prediction model group, respectively verifying the receiver operating characteristic curve (i.e., ROC curve) of all the training set plus the validation set data and calculating the area under the curve AUC value;

[0027] S62, ranking the prediction models according to the prediction accuracy of the positive samples in the verification results;

[0028] S63, select 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models:

[0029] The prediction model contains three parameters - project independent variables, coefficients of project independent variables and constants; during optimization, all project independent variables in all models are combined as the project independent variables of the final HCC risk prediction model; the coefficients of each project independent variable in all models are added and averaged as the coefficients of the corresponding project independent variables of the HCC risk prediction model; the constants in all models are added and averaged as the constants of the HCC risk prediction model.

[0030] By adopting the above technical solutions, the accuracy of HCC model prediction can be further improved.

[0031] Preferably, other physical examination index items as independent variables are screened by the following method:

[0032] First, verify the close correlation between the quantified transformation project and the endpoint event (through significance analysis, correlation analysis and rank sum analysis); if the significance of the correlation is P < 0.1, the project will be included in the candidate projects; otherwise, it will not be included in the candidate projects;

[0033] Secondly, for the selected items obtained by screening, the correlation between the items is checked (Box-Tidwell method, inter-item correlation analysis method and multicollinearity test method), and the items that meet the three assumptions of independent variables are selected to be included in the regression project, and finally used as independent variables together with the serum cytoplasmic thymidine kinase 1 test results; wherein, the three assumptions of independent variables are met, namely, there is a linear relationship between the item and the logit transformed value of the endpoint event, and there is no multicollinearity and significant correlation between the items.

[0034] By adopting the above technical solutions (especially for example, significance analysis, correlation analysis and rank sum analysis), it is possible to ensure that the test items have a significant or relatively significant correlation with the occurrence of the endpoint event, (for example, through the Box-Tidwell method), there is a linear relationship between the logit transformation values ​​of the continuous independent variables and the dependent variables, and (for example, through the inter-item correlation analysis method and the multicollinearity test method) the independence of the selected variable items is ensured. Including the independent variables that meet the above conditions in the "regression project" ensures that the univariate analysis difference of the independent variables is statistically significant, there is a linear relationship hypothesis between the continuous independent variables and the logit transformation values ​​of the dependent variables, and at the same time avoids the amplification and distortion of the influence weight of certain factors due to the inclusion of independent variable items with high correlation.

[0035] Preferably, after collecting the sample data, the following steps are also included: extracting the LI-RADS grading quantitative results in the text results of medical imaging (such as CT / MRI / B-ultrasound / CEUS) examinations, and performing natural number assignment processing on the classification option results, and converting the initial detection survey result data in different formats into consistent quantitative data.

[0036] By adopting the above technical solution, the medical imaging examination results and the initial different formats of the test survey results data are converted into consistent quantitative data, so as to facilitate the use in modeling. Further preferably, the LI-RADS grading quantitative results in the text results of the medical imaging examination are extracted using an endpoint event extraction tool, which specifically includes the following steps:

[0037] a. Output the number of occurrences of "LI" in the radiology text results through LEN and SUBSTITUTE commands;

[0038] b. Use the FIND command to obtain the location value N1 where “LI” first appears in the text, and use the MID command to grab the graded string after LI-RADS at position N1;

[0039] c. Take the first occurrence location value N1+1 as the starting position, continue to use the FIND command to obtain the second occurrence location value N2 of "LI", and grab the LI-RADS classification string at position N2;

[0040] d. Use the second occurrence of the location value N2+1 as the starting position and continue to use the FIND command to locate, and so on, to obtain the LI-RADS classification character string corresponding to all positions;

[0041] e. Use the VALUE and IFERROR commands to convert all the strings obtained at all positions into numbers (if there is no number, assign a value of "0");

[0042] f, use the MAX command to output the maximum value L of the converted number;

[0043] g. Use the IF command to convert the obtained L value into a risk rating number of 1 or 2.

[0044] By adopting the above technical solution, the endpoint event extraction tool can be used to specifically find LI-RADS classification information in massive medical imaging descriptive texts, and convert and output it according to the maximum value of its classification results, thereby improving conversion efficiency and reducing error rate.

[0045] Preferably, the classification option results are processed by assigning natural numbers, that is, a logical value conversion tool is used to convert the items into quantitative values ​​according to binary, ordered and unordered categories (for example, a negative result of a binary classification item is assigned to 1, and a positive result is assigned to 2; multi-classification items are converted into natural number series in sequence (such as 1, 2, to n)).

[0046] By adopting the above technical solution, the items are quantified according to the binary classification, ordered and disordered, which is conducive to using the physical examination indicator data for modeling.

[0047] In the aforementioned disease risk prediction modeling method based on LI-RADS grading, the HCC risk grouping information of the examinee is output by segmenting the risk probability results (such as segmenting according to CUTOFF = 0.500, those less than 0.5 are low-risk groups, and those greater than or equal to 0.5 are high-risk groups).

[0048] By adopting the above method, the examinees are divided into high-risk and low-risk groups, which is conducive to allocating most medical resources to the high-risk group.

[0049] In summary, the present application includes at least one of the following beneficial technical effects:

[0050] 1. This application sets the intermediate state of HCC that can be obtained in the physical examination sample (the intermediate state of HCC is an intermediate state that must be passed in the pathological process of HCC and has a certain level of deterioration risk characteristics) as the endpoint event (that is, the set of imaging examination results with a certain HCC risk is set as the risk prediction endpoint event (LR-3, LR-4 and LR-5 are set as positive results)), and multiple physical examination indicator items as independent variables. Since both the independent variables and the dependent variables are set in the physical examination stage, the risk prediction model can be developed by providing only the physical examination results, and there is no need to conduct a large number of inefficient follow-up work to collect clinical results, thereby reducing the cost and difficulty of obtaining positive result samples, and the stable accumulation of positive samples becomes possible. In the same time period, the iterative upgrade feasibility of the risk prediction model obtained by the system is higher. In addition, by combining multiple physical examination indicator items, mainly serum cytoplasmic thymidine kinase 1, i.e., TK1 test results, as independent variables, the prediction model obtained by modeling can be better applicable to the prediction of HCC risk; and in this application, multiple physical examination indicator items are used as independent variables, so the established HCC risk prediction model can be generally applicable to the general healthy population, such as being used in physical examination institutions to predict the HCC risk of ordinary physical examination populations, and can adapt to the high frequency and high throughput requirements of general screening.

[0051] 2. This application uses the unbalanced regression ILKL algorithm for modeling, which can effectively solve the problem of sample imbalance between positive samples and healthy samples in real physical examination data. The samples with the endpoint event of "general risk" are selected as the "majority group", and the positive sample group (i.e., LI-RADS grade of LR-3 and above) is the minority group. At this time, as long as the ratio of different states of the endpoint event (positive group / majority group) is randomly sampled in K categories, a "general risk" sample set equal to the number of samples in the positive group can be obtained; binary logistic regression is performed under the condition of sample equality, which can effectively avoid the problem that the prediction model is biased towards satisfying the majority group results due to sample imbalance. On the other hand, it also expands the diversity of the prediction model through the actual proportion of the positive population. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 is a method flow chart of an embodiment of the present application;

[0053] Figure 2 is a flow chart of the specific modeling method of the HCC risk prediction model in this application;

[0054] Figure 3 It is a flow chart of the method of project screening and modeling using the ILKL algorithm in this application;

[0055] Figure 4 A flow chart of the method for extracting LI-RADS grading quantitative results from the text results of medical imaging examinations using the endpoint event extraction tool;

[0056] Figure 5 A schematic diagram of a block diagram for evaluating the risk of HCC using the model established in this application;

[0057] Figure 6 Schematic diagram of the ROC curve for verifying the established prediction model in the experimental example. DETAILED DESCRIPTION

[0058] The following is combined with Figure 1-5 This application is described in further detail.

[0059] In the existing technology, for example, the aMAP risk scoring system proposed by Hou Jinlin et al. in 2020 only assesses the HCC risk of patients with chronic hepatitis. The HCC risk assessment system is modeled based on the patient's eventual hepatocellular carcinoma as the endpoint event, so it is necessary to take hepatocellular carcinoma patients as the target of sample collection. Since the number of confirmed liver cancer patients is less than one percent of the total population, the sample acquisition cost is high, the difficulty is high, and the sample accumulation efficiency is low, resulting in the iterative upgrade of the existing HCC risk assessment model. As the core of the risk assessment system, the iterative upgrade of the clinical prediction model is the key way to improve the effectiveness of the risk assessment system. However, because the physical examination data and clinical data are generally isolated into "information islands", the integrity of the physical examination data, the consistency of clinical results, and the correlation between the physical examination results and the clinical results (the time interval is too large, resulting in a lack of correspondence) will interfere with the effectiveness of the positive samples, resulting in the number of available positive samples being far lower than the actual situation. This increases the difficulty of predictive model research, and it is not easy to obtain an HCC risk prediction model with application value, and the feasibility of iterative risk prediction algorithms is not high. Based on this problem, the inventor conducted research and found that according to the latest version of LI-RADSv2018, liver nodules found by imaging can be divided into five categories: benign (LR-1), likely benign (LR-2), suspected HCC (LR-3), likely HCC (LR-4), and HCC (LR-5). According to the literature data of the 2014 and 2017 versions of LI-RADS disclosed in the CT / MRI LI-RADS v2018 CORE file, the positive predictive value (PPV) of HCC and malignant tumors is 0% at LR-1; the positive predictive value (PPV) of HCC is 16% at LR-2, and the PPV of malignant tumors is 18%; the PPV of HCC reaches 37% at LR-3, and the PPV of malignant tumors is 39%. The PPV values ​​of LR-4 and LR-5 are higher than those of LR-3, so the results of LR-1 and LR-2 can be regarded as the general risk status of HCC, and the results of LR-3, LR-4 and LR-5 can be regarded as the abnormal risk status of HCC.

[0060] Therefore, the inventors creatively came up with the idea of ​​setting the intermediate state of HCC that can be obtained in the physical examination sample (the intermediate state of HCC is an intermediate state that must be passed in the pathological process of HCC and has a certain level of deterioration risk characteristics) as the endpoint event (that is, converting the medical imaging examination results into a quantitative endpoint event state, setting the LI-RADS grade of LR-3 and above as the endpoint event (that is, setting LR-3, LR-4 and LR-5 as positive results)), and multiple physical examination indicator items as independent variables. Since both the independent variables and the dependent variables are set in the physical examination stage, it is only necessary to provide the physical examination results to develop the risk prediction model, and there is no need to conduct a large number of inefficient follow-up work to collect clinical results, which reduces the cost and difficulty of obtaining positive result samples, and makes it possible to accumulate positive samples stably. In the same time period, the iterative upgrade of the risk prediction model obtained by the system is more feasible. In addition, by combining multiple physical examination index items, mainly serum cytoplasmic thymidine kinase 1, or TK1 test results, as independent variables, the prediction model obtained by modeling can be better applicable to the prediction of HCC risk; and in this application, multiple physical examination index items are used as independent variables, so the HCC risk prediction model established can be generally applicable to the general healthy population, such as being used in physical examination institutions to predict the HCC risk of ordinary physical examination populations, and can adapt to the high frequency and high throughput requirements of general screening. That is, this application can effectively improve the development efficiency of the HCC risk prediction model by changing the data acquisition mode and reducing the acquisition cost of positive samples. On the basis of long-term and stable accumulation of positive samples, the HCC risk prediction model can perform more sufficient algorithm iterations to obtain a more stable and more predictively efficient HCC risk assessment system, and ultimately achieve efficient allocation of medical resources to high-risk HCC populations.

[0061] The present application embodiment discloses a disease risk prediction modeling method based on LI-RADS classification. Figure 1 , a disease risk prediction modeling method based on LI-RADS grading, including:

[0062] Step 1, collect sample data;

[0063] Step 2: convert the medical imaging examination results into quantitative endpoint event status, and set the LI-RADS grade of LR-3 and above as the endpoint event; use multiple physical examination index items, mainly serum cytoplasmic thymidine kinase 1 (TK1) test results, as independent variables;

[0064] Step 3: Modeling is performed using regression methods (such as the general formula of the binary logistic regression risk model) to obtain a risk prediction model for hepatocellular carcinoma (HCC).

[0065] In specific implementation, firstly, based on the overall observation of the physical examination results, the test results under the same testing system of the same institution can be selected, and then the different forms of the original items can be converted into a unified quantitative form through the endpoint event extraction tool, logical value conversion tool, etc.; then, through the data cleaning step, the samples with data integrity and standardization defects are removed (such as removing samples with no records of TK1 values ​​or LI-RADS results, and removing samples whose quantitative results after conversion are obviously beyond the measurement range, such as negative numbers or outlier samples obtained during the test (the judgment of outliers is based on the method of 6.2.1 "upper side situation" in "Statistical Processing and Interpretation of Data Normal Sample Outliers Judgment and Processing" in GBT 4883-2008)). Finally, the unstructured database can be used to store the original state and the project results of the numerical conversion, and output them into analyzable data file format according to the analysis requirements.

[0066] Generally, physical examination result items are divided into three categories according to the form of item results: 1) medical imaging examination results, including ultrasound / contrast ultrasound (CEUS) items, CT items and MRI items; 2) continuous value items; 3) classification result items, including binary classification items, ordered multi-classification items and unordered multi-classification items.

[0067] Optionally, the model is modeled by regression method to obtain a risk prediction model for hepatocellular carcinoma HCC, such as Figure 2 As shown, the specific steps include:

[0068] S1, use three-fold cross-validation method to randomly sample and split the original sample data into training set and validation set;

[0069] S2, for different training sets, multiple prediction models are obtained by repeatedly using the unbalanced regression ILKL algorithm; multiple prediction models obtained corresponding to the same training set form a prediction model library;

[0070] S3, screening the prediction models in each prediction model library, if the prediction model formula contains the TK1 project independent variable, then the prediction model is included in the initial screening prediction model group;

[0071] S4, for the prediction model in the initial screening prediction model group, the corresponding validation set samples were used to perform receiver operating characteristic curve (ROC curve) validation and calculate the area under the curve (AUC) value;

[0072] S5, if the AUC value is greater than or equal to 0.7, the prediction model in the corresponding preliminary screening prediction model group is included in the final prediction model group;

[0073] S6. Optimize each prediction model in the final prediction model group according to a parameter comprehensive method to obtain an HCC risk prediction model.

[0074] In specific implementation, for example, the number of samples used for model development is 10,396, which are divided into three groups through three-fold sampling: Group 1 is 3,465 samples, Group 2 is 3,465 samples, and Group 3 is 3,466 samples. The ILKL algorithm is used for sample matching in model development, and the ROC curve verification method is used for model screening. The three-fold method uses Group 1 as the validation set, and Groups 2 and 3 are combined to form the training set. Figure 3 , by screening the items in the training set and looping the ILKL algorithm, we obtain the prediction model library 1 (containing 100 prediction models). Then, we repeat the above steps in the combination 2 and combination 3 of the three-fold sampling method to form different prediction model libraries.

[0075] Optional, such as Figure 3 As shown, the modeling using the unbalanced regression ILKL algorithm described in step S2 specifically includes the following steps:

[0076] S21, samples were divided into two groups according to the endpoint event status, where samples with “average risk” (i.e., LR-1 and LR-2) as the “majority group” and positive samples (i.e., LI-RADS grade LR-3 and above) as the minority group;

[0077] S22, setting the physical examination index item as the independent variable as the clustering index variable, setting the endpoint event as the sample label, and setting K (a natural number greater than 2 and less than or equal to 10) as the number of category groups;

[0078] S23, randomly selecting K samples from the multi-array samples as centroids, for each point in the multi-array samples, calculating the Euclidean distance between it and each centroid sample based on the clustering indicator variable, and dividing it into the set to which the centroid with the shortest distance belongs;

[0079] S24, after all data are classified into K sets, the centroid of each set is recalculated;

[0080] S25, if the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold (i.e., convergence is reached), the obtained K sets are used as K-Means clustering groups; otherwise, the iteration is continued until the K-Means clustering groups are finally obtained;

[0081] S26, random sampling of K-Means clustering groups according to the ratio of minority group to majority group;

[0082] S27, merge all the sampled samples with the minority group samples, and obtain the prediction model through binary logistic regression.

[0083] For example, if there are 120 positive samples, in the stage of sampling the healthy group, random sampling is performed on the K-Means clustering grouping according to the positive group / healthy group ratio, so that the number of samples after sampling the healthy group is close to 120, thereby achieving the purpose of reducing the prediction bias.

[0084] Optionally, the model optimization according to the parameter comprehensive method described in step S6 specifically includes the following steps:

[0085] S61, using each prediction model in the final prediction model group, respectively verifying the receiver operating characteristic curve (i.e., ROC curve) of all the training set plus the validation set data and calculating the area under the curve AUC value;

[0086] S62, ranking the prediction models according to the prediction accuracy of the positive samples in the verification results;

[0087] S63, select 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models:

[0088] The prediction model contains three parameters - project independent variables, coefficients of project independent variables and constants; during optimization, all project independent variables in all models are combined as the project independent variables of the final HCC risk prediction model; the coefficients of each project independent variable in all models are added and averaged as the coefficients of the corresponding project independent variables of the HCC risk prediction model; the constants in all models are added and averaged as the constants of the HCC risk prediction model.

[0089] For example, when performing parameter optimization, the prediction model contains three parameters: A n A1 is a fixed number for all independent variable items in the model. n ; B m-n is the item coefficient, which refers to the A in model M (for example, M=1~5). n Parameters of the project, actual number is B 1-1 To B 5-n , if any model does not contain a certain A n The project parameter B is set to 0; C M is the constant of model M, and the actual number is C1-C5. The optimized model contains all the A in the five merged models. n The optimization of parameters B and C is as follows:

[0090] C op = (C1+C2+C3+C4+C5) / 5

[0091] B op(m-n) =(B 1-n +B 2-n +B3-n +B 4-n +B 5-n ) / 5.

[0092] Optional, such as Figure 3 As shown, other physical examination index items as independent variables were screened by the following method:

[0093] First, verify the close correlation between the quantified transformation project and the endpoint event (through significance analysis, correlation analysis and rank sum analysis); if the significance of the correlation is P < 0.1, the project will be included in the candidate projects; otherwise, it will not be included in the candidate projects;

[0094] Secondly, for the candidate items obtained by screening, the correlation between the items was checked (Box-Tidwell method, inter-item correlation analysis method and multicollinearity test method), and the items that met the three assumptions of the independent variables were selected to be included in the regression project, and finally used as the independent variable together with the serum cytoplasmic thymidine kinase 1 test results; among them, the three assumptions of the independent variables were met, namely, there was a linear relationship between the item and the endpoint event (logit) conversion value, and there was no multicollinearity and significant correlation between the items.

[0095] When collecting samples, the following information of the examinee is collected: 1) basic information of the examinee (sensitive information removed); 2) family medical history / personal medical history; 3) lifestyle survey results; 4) single test results; 5) medical imaging examination results. Then the physical examination items and the converted quantitative items are imported into the formed database storage system. According to the logical relationship of the recorded content, it is divided into basic information of the examinee, family medical history / personal medical history, lifestyle survey results, single test results and medical imaging examination results. Among them, the single test results are composed of blood tests, biochemical tests, infection immunity tests and body fluid tests.

[0096] Optionally, after collecting the sample data, the following steps are also included: extracting the LI-RADS grading quantitative results from the text results of medical imaging (such as CT / MRI / B-ultrasound / CEUS) examinations, and performing natural number assignment processing on the classification option results, so as to convert the initial detection survey result data in different formats into consistent quantitative data.

[0097] Optionally, the endpoint event extraction tool is used to extract the LI-RADS grading quantitative results from the text results of medical imaging examinations, such as Figure 4 As shown, the specific steps include:

[0098] a. Output the number of occurrences of "LI" in the radiology text results through LEN and SUBSTITUTE commands;

[0099] b. Use the FIND command to obtain the location value N1 where “LI” first appears in the text, and use the MID command to grab the graded string after LI-RADS at position N1;

[0100] c. Take the first occurrence location value N1+1 as the starting position, continue to use the FIND command to obtain the second occurrence location value N2 of "LI", and grab the LI-RADS classification string at position N2;

[0101] d. Use the second occurrence of the location value N2+1 as the starting position and continue to use the FIND command to locate, and so on, to obtain the LI-RADS classification character string corresponding to all positions;

[0102] e. Use the VALUE and IFERROR commands to convert all the strings obtained at all positions into numbers (if there is no number, assign a value of "0");

[0103] f, use the MAX command to output the maximum value L of the converted number;

[0104] g. Use the IF command to convert the obtained L value into a risk rating number of 1 or 2.

[0105] In specific implementation, for example, the command line in EXCEL 2013 is used to extract the LI-RADS grading quantitative results from the text results of medical imaging examinations:

[0106] Step 1):

[0107] ƒx=0.5*(LEN(initial text)-LEN(SUBSTITUTE(initial text,"LI","")));

[0108] Step 2):

[0109] ƒx=MID(A3,FIND("LI",initial text)+8,1);

[0110] Step 3):

[0111] ƒx=MID(A3,FIND("LI",initial text,N1+1)+8,1);

[0112] Step 4):

[0113] ƒx=MID(A3,FIND("LI",initial text,N2+1)+8,1);

[0114] Step 5):

[0115] ƒx=MID(A3,FIND("LI",text table,N3+1)+8,1);

[0116] Step 6):

[0117] ƒx=IFERROR(VALUE(capture LI-RADS grade string),0);

[0118] Step 7):

[0119] ƒx=MAX(Step 2 extracts the classification value: Step 5 extracts the classification value);

[0120] Step 8):

[0121] ƒx==IF(L>2,2,1).

[0122] Optionally, the classification option results are processed by assigning natural numbers, that is, a logical value conversion tool is used to convert the items into quantitative values ​​according to binary, ordered and unordered categories (for example, binary classification items are converted to 0 / 1 or 1 / 2 according to positive / negative results, such as male / female can be converted to 1 / 2 respectively; LR-1 and LR-2 are converted to 1, LR-3, LR-4 and LR-5 are uniformly converted to 2; multi-classification items are converted into natural number series in sequence (such as 1, 2, to n)).

[0123] In the present application, the HCC risk grouping information of the examinee can be output by segmenting the risk probability results (such as segmenting according to CUTOFF = 0.500, those less than 0.5 are low-risk groups, and those greater than or equal to 0.5 are high-risk groups).

[0124] In specific implementation, long-term physical examination data from the same location and institution can be collected, including TK1 testing and other routine health examination items (including but not limited to blood routine, liver function, tumor markers, etc.). Figure 5 As shown, according to the HCC risk prediction model established in this application, select specific items in the physical examination report; import the prediction model in combination with the test results of the TK1 kit; calculate and judge the risk probability through the model; and output the text of the prediction result according to the calculation result. The HCC risk prediction model established in this application can be applied to ordinary healthy people, and can be specifically used in physical examination institutions to predict the risk of HCC.

[0125] Experimental example:

[0126] The physical examination data of a PLA hospital in 2017 (total number of samples: 10,639), including 130 people in the positive group and 10,509 people in the negative group (in line with the actual proportion, improving the extensibility of the regression model of this application and the practical significance of the risk assessment method), were used to model the model based on the modeling method of this application. The final prediction model used age classification and 11 biomarkers or blood indicators as independent variables: cytoplasmic thymidine kinase 1 concentration (TK1), lymphocyte count, mean red blood cell volume, platelet volume, albumin ratio, serum albumin, aspartate aminotransferase, creatinine, urea creatinine, alpha-fetoprotein (AFP) and carcinoembryonic antigen (CEA).

[0127] Among them, the age level value can be preset as follows: 20-29 years old is assigned to 1, 30-39 years old is assigned to 2, 40-49 years old is assigned to 3, 50-59 years old is assigned to 4, 60-69 years old is assigned to 5, 70-79 years old is assigned to 6, and 80 years old and above is assigned to 7; TK1 value (pM) detection can be carried out using the CIS series chemiluminescence digital imaging analyzer and thymidine kinase 1 diagnostic kit produced by Shenzhen Huarui Tongkang Biotechnology Co., Ltd.; lymphocyte count, mean red blood cell volume, platelet volume, albumin ratio, and serum albumin can be detected using the XFA6100 fully automatic blood cell analyzer; aspartate aminotransferase, creatinine, and urea creatinine can be detected using the MD-100 semi-automatic biochemical analyzer; alpha-fetoprotein (AFP) and carcinoembryonic antigen (CEA) can be measured using the ELISA detection kit (supplier: Shanghai Guandao Bioengineering Co., Ltd.) and the fully automatic biochemical detector.

[0128] The function formula of the corresponding hepatocellular carcinoma (HCC) risk prediction model is: risk probability P = ExpX / (1+ExpX), where X = 8.039744953 + (0.2317891 × age level) + (0.5909724 × TK1) + (0.0386128 × lymphocyte count) + (- 0.016474096 × average red blood cell volume) + (-8.131866143 × platelet volume) + (-0.215379 × albumin ratio) + (-0.151802523 × serum albumin) + (-0.021515429 × aspartate aminotransferase) + (0.017035858 × creatinine) + (-3.000811143 × urea creatinine) + (0.01944881 × AFP) + (-0.004730143 × CEA).

[0129] The risk probability P is the possibility of LI-RADS 3 results in medical imaging, that is, the LI-RADS 3 classification results in the liver imaging report and data system are used as the endpoint event of the prediction model, anchoring the determined risk probability of hepatocellular carcinoma; when the risk probability P <0.5, it can be judged as a general risk; and when the risk probability P ≥0.5, it can be judged as an abnormal risk.

[0130] In addition, the physical examination data of a PLA hospital in 2017 (total number of samples: 10,639) were used to verify the prediction model. Figure 6 As shown, the AUC value was 0.730, 95% CI 0.722-0.739, Youden index was 0.3633, sensitivity was 73.08, and specificity was 63.25 (threshold 0.5).

[0131] In addition, the inventors also compared the value of combined detection methods. The results showed that the simultaneous detection of 12 items in this experimental example had a higher prediction accuracy of the HCC risk prediction model finally established compared with the use of markers TK1 (<2 is normal), AFP (<35 is normal), CEA (<5 is normal) alone or the combined use of TK1, AFP, and CEA.

[0132] The above are all preferred embodiments of the present application, and the protection scope of the present application is not limited thereto. Therefore, all equivalent changes made according to the methods and principles of the present application should be included in the protection scope of the present application.

Claims

1. A disease risk prediction modeling method based on LI-RADS grading, characterized in that: include: Collect sample data; The results of medical imaging examinations were converted into quantitative endpoint events, and the LI-RADS grade of LR-3 and above was set as the endpoint event; A variety of physical examination indicators, mainly the results of serum cytoplasmic thymidine kinase 1 (TK1), were used as independent variables; a model was built using regression method to obtain a risk prediction model for hepatocellular carcinoma (HCC); The model is built by regression method to obtain a risk prediction model for hepatocellular carcinoma (HCC). The following steps are involved: S1, use three-fold cross-validation method to randomly sample and split the original sample data into training set and validation set; S2, for different training sets, multiple prediction models are obtained by repeatedly using the unbalanced regression ILKL algorithm; multiple prediction models obtained corresponding to the same training set form a prediction model library; S3, screening the prediction models in each prediction model library, if the prediction model formula contains the TK1 project independent variable, then the prediction model is included in the initial screening prediction model group; S4, for the prediction model in the initial screening prediction model group, the corresponding validation set samples were used to perform receiver operating characteristic curve validation and calculate the area under the curve AUC value; S5, if the AUC value is greater than or equal to 0.7, the prediction model in the corresponding preliminary screening prediction model group is included in the final prediction model group; S6, optimizing each prediction model in the final prediction model group according to a parameter comprehensive method, thereby obtaining an HCC risk prediction model; The model optimization according to the parameter synthesis method described in step S6 specifically includes the following steps: S61, using each prediction model in the final prediction model group, respectively performing receiver operating characteristic curve verification on the data of all training sets plus the validation set and calculating the area under the curve AUC value; S62, ranking the prediction models according to the prediction accuracy of the positive samples in the verification results; S63, select 2-10 models with the highest prediction accuracy and optimize the parameters of the prediction models: The prediction model contains three parameters - project independent variables, coefficients of project independent variables and constants; during optimization, all project independent variables in all models are combined as the project independent variables of the final HCC risk prediction model; the coefficients of each project independent variable in all models are added and averaged as the coefficients of the corresponding project independent variables of the HCC risk prediction model; the constants in all models are added and averaged as the constants of the HCC risk prediction model.

2. The disease risk prediction modeling method based on LI-RADS grading according to claim 1, characterized in that: The modeling using the unbalanced regression ILKL algorithm described in step S2 specifically includes the following steps: S21, the samples are divided into two groups according to the endpoint event status, where the samples with the endpoint event of "general risk" are the "majority group" and the positive sample group is the minority group; S22, setting the physical examination index item as the independent variable as the clustering index variable, setting the endpoint event as the sample label, and setting K as the number of category groups; S23, randomly selecting K samples from the multi-array samples as centroids, for each point in the multi-array samples, calculating the Euclidean distance between it and each centroid sample based on the clustering indicator variable, and dividing it into the set to which the centroid with the shortest distance belongs; S24, after all data are classified into K sets, the centroid of each set is recalculated; S25, if the Euclidean distance between the newly calculated centroid sample and the original centroid sample is less than the set threshold, the obtained K sets are used as K-Means clustering groups; otherwise, the iteration is continued until the K-Means clustering groups are finally obtained; S26, random sampling of K-Means clustering groups according to the ratio of minority group to majority group; S27, merge all the sampled samples with the minority group samples, and obtain the prediction model through binary logistic regression.

3. The disease risk prediction modeling method based on LI-RADS grading according to claim 1, characterized in that: Other physical examination index items as independent variables were screened by the following method: First, verify the close correlation between the quantified transformation project and the endpoint event; if the significance of the correlation is P<0.1, the project will be included in the candidate projects; otherwise, it will not be included in the candidate projects; Secondly, for the candidate items obtained by screening, the correlation between the items was checked, and the items that met the three assumptions of the independent variables were selected and included in the regression items, and finally used as independent variables together with the serum cytoplasmic thymidine kinase 1 test results; among them, the three assumptions of the independent variables were met, namely, there was a linear relationship between the item and the logit transformed value of the endpoint event, and there was no multicollinearity and significant correlation between the items.

4. The disease risk prediction modeling method based on LI-RADS grading according to claim 1, characterized in that: After collecting sample data, the following steps are also included: extracting the LI-RADS grading quantitative results in the text results of medical imaging examinations, and performing natural number assignment processing on the classification option results, and converting the initial test survey result data in different formats into consistent quantitative data.

5. The disease risk prediction modeling method based on LI-RADS grading according to claim 4 is characterized in that: The endpoint event extraction tool was used to extract the LI-RADS grading quantitative results from the text results of medical imaging examinations. The following steps are involved: a. Output the number of occurrences of "LI" in the radiology text results through LEN and SUBSTITUTE commands; b. Use the FIND command to obtain the location value N1 where "LI" first appears in the text, and use the MID command to grab the graded string after LI-RADS at position N1; c. Take the first occurrence location value N1+1 as the starting position, continue to use the FIND command to obtain the second occurrence location value N2 of "LI", and grab the LI-RADS classification string at position N2; d. Use the second occurrence of the location value N2+1 as the starting position and continue to use the FIND command to locate, and so on, to obtain the LI-RADS classification character string corresponding to all positions; e. Use the VALUE and IFERROR commands to convert all the strings obtained at all positions into numbers; f, use the MAX command to output the maximum value L of the converted number; g. Use the IF command to convert the obtained L value into a risk rating number of 1 or 2.

6. The disease risk prediction modeling method based on LI-RADS grading according to claim 4, characterized in that: The classification option results are processed by assigning natural numbers, that is, the items are converted into quantitative values ​​according to binary classification, ordered and unordered by using a logical value conversion tool.

7. The disease risk prediction modeling method based on LI-RADS grading according to claim 1, characterized in that: By segmenting the risk probability results, the HCC risk grouping information of the subject is output.

Citation Information

Patent Citations

  • Early diagnosis method and diagnostic kit for liver cancer

    CN109557312A

  • Multi-physiological data fusion analysis method based on neural network

    CN110378394A

  • Disease suffering risk prediction method and device

    CN110838366A