A method for evaluating the efficacy of conversion therapy for liver cancer

By extracting significant image features from the CT images of liver cancer patients and using unsupervised system clustering and random forest models to evaluate the efficacy of liver cancer transformation treatment, the problem of difficulty in accurately evaluating treatment effects in the prior art is solved, and higher evaluation accuracy and interpretability are achieved.

CN119672035BActive Publication Date: 2025-06-06ANHUI PROVINCIAL HOSPITAL +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510204723.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-06
Estimated Expiration
2045-02-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively analyze and evaluate CT image data of liver cancer patients, and it is impossible to accurately determine the success rate and subtype of translational treatment, which makes it difficult to determine the treatment plan.

Method used

By extracting significant image features from the CT images of liver cancer patients, subclass divisions were performed using unsupervised system clustering, and supervised learning was combined with random forest models to predict the efficacy of liver cancer transformation treatment.

Benefits of technology

It improves the accuracy and interpretability of the efficacy evaluation of liver cancer conversion treatment, can more effectively evaluate the treatment effect and adaptability of patients, and optimize the treatment plan.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119672035B_ABST
    Figure CN119672035B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for implementing the evaluation of the therapeutic effect of liver cancer conversion therapy, comprising: in the computed tomography (CT) images of liver cancer patients, using pre-selected image features with significant differences as the basis for judging the Euclidean distance in sub-classification, using unsupervised system clustering to divide the CT images into various sub-class clusters as unlabeled samples, and defining a label for each sub-class cluster; using the various samples and labels contained in the unlabeled samples as training sets, and using supervised learning to train a random forest model. The implementation of the present invention combines unsupervised with supervised learning, first using unsupervised system clustering to label the unlabeled samples, and then using supervised learning to train the random forest model, thereby effectively improving the accuracy and interpretability of the prediction of the efficacy evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent data analysis and processing, and in particular to a method for implementing efficacy evaluation of liver cancer conversion therapy. Background Art

[0002] Hepatocellular carcinoma (HCC) is one of the malignant tumors with the highest mortality rate worldwide. Traditional treatments have limited effects on patients with advanced HCC, while emerging immunotherapy combined with targeted drug conversion therapy provides new treatment opportunities for these patients. However, due to the heterogeneity of tumors, the efficacy of advanced unresectable HCC patients receiving immunotherapy combined with targeted conversion therapy varies. Conversion therapy is an important way to improve the survival of patients with advanced liver cancer, which is to convert unresectable liver cancer into resectable liver cancer and then remove the tumor.

[0003] High-dimensional radiomics features in CT images of HCC patients before treatment can reflect the heterogeneity of tumors, which is related to the efficacy of conversion therapy. However, traditional CT analysis is mostly limited to qualitative diagnosis and simple quantification, and cannot fully explore the potential information in the images. In addition, the information in the images has high data feature dimensions and a large number of complex features, which makes it difficult to accurately determine the success rate of conversion therapy for liver cancer patients and their subtypes. Therefore, it is of great significance to analyze the classification rules of medical imaging data of liver cancer patients and build an effective classification model.

[0004] At present, classification methods based on machine learning have been successfully applied in many fields of classification with their powerful feature representation capabilities and end-to-end learning mechanisms. Among them, the KNN classification algorithm based on distance measurement takes the entire data set as the training set, determines the distance between the sample to be classified and each training sample, and then finds the K samples closest to the sample to be classified as the K nearest neighbors of the sample to be classified. The sample category to be classified is the category with the largest proportion. The early decision tree (DT) classification algorithm, the CART algorithm, uses a tree structure algorithm to divide data into discrete categories. The DT classification algorithm generally has two steps: one is to use the training set to start from the top-level root node of the DT, and judge from top to bottom to form a decision tree (i.e., establish a classification model); the second is to use the built DT to classify the sample set to be classified. However, the flexible classification model based on DT has the disadvantages of being prone to overfitting and ignoring the correlation of attributes in the data set. Another support vector machine (SVM) classification algorithm is mainly used in the field of binary classification. In addition, previous unsupervised learning methods have also been widely used in classification tasks, and representative methods include Kmeans, system clustering, and hierarchical clustering. This type of method classifies data based on the numerical characteristics of the data itself. However, due to the lack of supervised information, the interpretability of the classification basis is often not clear enough, making it difficult to provide clear theoretical support for the classification results.

[0005] The above algorithms are all single classification methods, and in actual applications, there are still problems that these algorithms cannot effectively solve, that is, the above algorithms cannot effectively analyze and process the CT image data of HCC patients. This means that the current analysis of CT images of HCC patients can only reflect the similarity of pathological features in the samples, but cannot provide the basis for the corresponding subclassification, that is, it is impossible to clearly explain the basis for classification, and thus it is impossible to effectively evaluate and analyze the conversion treatment effect of HCC patients, which brings great trouble to the determination of subsequent treatment plans.

[0006] In view of this, the present invention is proposed. Summary of the invention

[0007] The purpose of the present invention is to provide a method for evaluating the efficacy of liver cancer conversion therapy, so as to effectively evaluate and analyze the conversion therapy effect on HCC patients and solve the problems existing in the prior art.

[0008] The objective of the present invention is achieved through the following technical solutions:

[0009] A method for evaluating the efficacy of liver cancer conversion therapy, comprising:

[0010] In the computer tomography CT images of liver cancer patients, image features with significant differences pre-selected based on significant image features are used as the basis for judging the Euclidean distance in subclass division, and the CT images are divided into various subclass clusters as unlabeled samples by unsupervised system clustering; wherein, image features corresponding to surgical success rate-related factors in the CT images of liver cancer patients are collected and extracted, and the degree of deviation between the actual observed value and the theoretical inferred value of the image features as samples is statistically analyzed by chi-square test analysis, and then the significant image features are determined according to the statistical results;

[0011] A label is defined for each subclass cluster as a label value of each sample included in the unlabeled sample; wherein the subclasses include a subclass of successful surgery and a subclass of unresectable surgery;

[0012] Each sample included in the unlabeled sample and the label value determined for it are used as a training set, and a random forest model is trained by supervised learning; the random forest model is used to predict and evaluate the efficacy of liver cancer conversion therapy for patients with unknown types of liver cancer.

[0013] The method for obtaining the image features with significant differences pre-selected based on the significant image features includes:

[0014] Among the obtained imaging features of liver cancer patients, the corresponding significant imaging features are determined based on the factors related to the surgical success rate, and the standard deviation of each significant imaging feature is calculated;

[0015] Based on the standard deviation, using a box plot to capture outliers in the significant image features;

[0016] Based on the captured outliers in the significant image features, the image features with significant differences are selected and determined as image features for subclass division indicators.

[0017] The image features include: tumor shape, size, texture and density distribution, as well as liver volume, gray level co-occurrence matrix GLCM and local binary pattern.

[0018] The scanning parameters of the CT image include:

[0019] Voltage 80-120KV, current 100-300mA, rotation time 0.3-0.5s, detector combination 0.625×64mm, pitch 0.8-0.9, bed feed speed 30-50mm / s, layer thickness 3-7mm, spacing 3-5mm.

[0020] The processing of determining the significant image features according to the statistical results includes:

[0021] The image features and the success or failure of the operation were taken as two variables, and an interactive analysis was performed on the two variables based on the social science statistical software SPSS. According to the significance P value of each image feature obtained by the interactive analysis, it was determined whether the corresponding image feature was the significant image feature.

[0022] The standard deviation includes: overall standard deviation, sample standard deviation, standard error and sample mean, and the corresponding calculation formula is:

[0023] The population standard deviation is: ; The sample standard deviation is: ; standard error is: ; The sample mean is: ;

[0024] in, Represents the overall mean of an indicator. represents the index value of the i-th sample point, n represents the number of sample data, Represents the mean of the sample; the population standard deviation is often used to measure the degree of data dispersion of the entire population, and the sample standard deviation is used to measure the degree of dispersion of sample data.

[0025] The process of using a box plot to capture outliers in the significant image features includes:

[0026] In the box plot, the standard deviations are sorted from small to large; the lower quartile Q1 is the value ranked 25%, the upper quartile Q3 is the value ranked 75%, and the interquartile range ;

[0027] Based on the box plot, capturing and determining the abnormal value according to a predetermined interval parameter;

[0028] The interval parameters include:

[0029] ;

[0030] And the process of capturing and determining the abnormal value includes: identifying the image feature corresponding to the standard deviation outside the interval corresponding to the interval parameter as an abnormal value;

[0031] The process of selecting and determining the image features with significant differences includes:

[0032] The significant image features corresponding to the outliers are selected and removed from the significant image features, and the remaining other significant image features are selected as the image features with significant differences to be used as image features for subclass division indicators.

[0033] The subcategories of successful surgery include: high success rate and moderate efficacy; the subcategories of unresectable surgery include: inability to withstand surgical trauma, insufficient remaining liver volume, and poor efficacy after resection.

[0034] The process of training the random forest model using supervised learning includes:

[0035] Performing multiple random sampling of samples and features on the image features used as subclass division indicators, and dividing the area where the subsets obtained by each sampling are located;

[0036] Based on the area where the subset is located, the optimal feature number and split point value corresponding to the subset are found by traversing all samples and features; when the number of samples contained in the leaf nodes of the subset is less than a predetermined value, a decision tree DT model of the subset is obtained;

[0037] After the corresponding multiple DT models are constructed for all subsets obtained through multiple samplings, the random forest model is constructed based on the multiple DT models.

[0038] The process of obtaining the decision tree DT model of the subset includes:

[0039] The independent variables corresponding to the image features of the subclass division index are analyzed by using the bootstrap random sampling method with replacement and the RSM algorithm for randomly selecting feature subsets. conduct The samples and features are randomly sampled, and the Secondary generation subset The process is as follows:

[0040] ;

[0041] in, is a random subspace function, is a random sampling function; represents the number of features selected by the sub-training set, , where M is the number of features in the independent variable and N is the number of samples;

[0042] Divide the subset The area where it is located, and find the optimal feature number and split point value by traversing all its samples and features , the value for:

[0043] ;

[0044] in, and Indicated in and The measured value of an image feature of a region, and express and The average of the measurements in the area, is the sample number threshold contained in the leaf node, st refers to and The constraints that need to be met; represents the optimal variable number; Represents the value of the split point;

[0045] The sub-set The area and the corresponding output value are determined using the following criteria:

[0046] ;

[0047] For the area and Repeat the steps of finding the optimal feature number and split point value until the number of samples in the leaf node of the subset is less than the set threshold ; Accordingly, the input space is divided into The DT model of this subset is obtained as follows:

[0048]

[0049] ;

[0050] in, represents the decision tree model on the jth subset, Represents the number of divided areas; The area contains samples indivual; express In the region In the subset a truth value; represents an indicator function, when When present ,otherwise ; represents the number of features selected by the sub-training set, Indicates the candidate split point value under a certain feature;

[0051] And the process of constructing the random forest model based on multiple DT models includes:

[0052] exist Repeat the above process to obtain the decision tree DT model of this subset times, obtained After the DT model, the corresponding random forest sub-model is:

[0053] ;

[0054] in, Represents the integrated random forest model.

[0055] Compared with the prior art, in the implementation method of the efficacy evaluation of liver cancer conversion therapy provided by the present invention, the chi-square test can be used to analyze and determine the correlation between the surgical success rate and features such as tumor shape, texture, density distribution, liver size and gray-level co-occurrence matrix (GLCM), so as to provide guidance for subsequent decision-making; moreover, it also uses the standard deviation as a measurement standard to reflect the degree of discreteness of each image feature, so as to help select appropriate features as indicators for subclass division; and identifies and excludes outliers with small standard deviations through box plots, thereby making the subclass classification more reasonable; further, the present invention also combines unsupervised and supervised learning, first using unsupervised system clustering to annotate unlabeled samples, and then using supervised learning to train the random forest model, thereby effectively improving the accuracy and interpretability of efficacy evaluation predictions. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0057] Figure 1 A schematic diagram of a process for determining factors related to the success rate of conversion therapy provided by an embodiment of the present invention;

[0058] Figure 2 A schematic diagram of the conversion therapy effect evaluation process provided by an embodiment of the present invention;

[0059] Figure 3 Schematic diagram of sub-classification of conversion therapy effects provided in embodiments of the present invention Figure 1 ;

[0060] Figure 4 Schematic diagram of sub-classification of conversion therapy effects provided in embodiments of the present invention Figure 2 . DETAILED DESCRIPTION

[0061] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention; it is obvious that the described embodiments are only part of the embodiments of the present invention, not all of the embodiments, which does not constitute a limitation of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the protection scope of the present invention.

[0062] First, the terms that may be used in this article are explained as follows:

[0063] The term “and / or” means that either or both of them can be realized at the same time. For example, X and / or Y means both “X” or “Y” and “X and Y”.

[0064] The terms "include", "comprises", "contains", "has" or other descriptions with similar semantics should be interpreted as non-exclusive inclusion. For example, including certain technical feature elements (such as raw materials, components, ingredients, carriers, dosage forms, materials, dimensions, parts, components, mechanisms, devices, steps, procedures, methods, reaction conditions, processing conditions, parameters, algorithms, signals, data, products or products, etc.) should be interpreted as including not only certain technical feature elements explicitly listed, but also other technical feature elements known in the art that are not explicitly listed.

[0065] The term "consisting of..." means excluding any technical feature elements not explicitly listed. If this term is used in a claim, it will make the claim closed, so that it does not contain technical feature elements other than the technical feature elements explicitly listed, except for the conventional impurities related to them. If this term only appears in a clause of a claim, it only limits the elements explicitly listed in the clause, and the elements recorded in other clauses are not excluded from the overall claim.

[0066] The term "parts by mass" refers to the mass ratio relationship between multiple components. For example, if it is described that component X is x parts by mass and component Y is y parts by mass, then the mass ratio of component X to component Y is x:y. 1 part by mass can represent any mass, for example, 1 part by mass can be represented as 1 kg or 3.1415926 kg. The sum of the parts by mass of all components is not necessarily 100 parts, but can be greater than 100 parts, less than 100 parts or equal to 100 parts. Unless otherwise specified, the parts, proportions and percentages described herein are all measured by mass.

[0067] Unless otherwise specified or limited, the terms "installed", "connected", "connected", "fixed" and the like should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in this article can be understood according to specific circumstances.

[0068] When concentration, temperature, pressure, size or other parameters are expressed in the form of a numerical range, the numerical range should be understood to specifically disclose all ranges formed by the pairing of any upper limit, lower limit, and preferred value in the numerical range, regardless of whether the range is explicitly stated; for example, if a numerical range of "2 to 8" is stated, the numerical range should be interpreted as including ranges such as "2 to 7", "2 to 6", "5 to 7", "3 to 4 and 6 to 7", "3 to 5 and 7", "2 and 5 to 7", etc. Unless otherwise specified, the numerical ranges stated herein include both their end values ​​and all integers and fractions within the numerical range.

[0069] The orientation or position relationship indicated by terms such as "center", "longitudinal", "lateral", "length", "width", "thickness", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", "clockwise", "counterclockwise", etc. are based on the orientation or position relationship shown in the drawings and are only for the convenience and simplification of description, and do not explicitly or implicitly indicate that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation of this document.

[0070] The following is a detailed description of a method for evaluating the efficacy of liver cancer conversion therapy provided by the present invention. The contents not described in detail in the embodiments of the present invention belong to the prior art known to professional and technical personnel in the field. If no specific conditions are specified in the embodiments of the present invention, the conventional conditions in the field or the conditions recommended by the manufacturer are followed. The reagents or instruments used in the embodiments of the present invention, if the manufacturer is not specified, are all conventional products that can be purchased commercially.

[0071] The purpose of the present invention is to take HCC, one of the serious diseases that endangers the health of the whole people, as the object, to understand the factors related to its conversion treatment success rate, so as to assist in accurately judging whether the patient can accept liver resection and ensure the safety of the operation. The technical solution provided by the present invention mainly includes the integration of CT image data acquisition, feature extraction, feature screening based on chi-square test and standard deviation, unsupervised system clustering annotation and random forest model prediction and other processing links, thereby forming a complete liver cancer conversion treatment efficacy evaluation process. Each link is interrelated and influences each other, and together achieves the goals of personalized evaluation of surgical success rate. For example, the parameter settings of CT image data acquisition (such as voltage, current, rotation time, etc.) affect the accuracy of subsequent feature extraction, and the results of feature screening are directly related to the input data quality of system clustering and random forest models, which in turn affects the accuracy of the final efficacy evaluation.

[0072] Furthermore, the embodiment of the present invention provides a method for evaluating the efficacy of liver cancer conversion therapy, that is, an implementation scheme for interpretable liver cancer conversion therapy efficacy evaluation based on system clustering and random forest (RF). The specific implementation process of the implementation scheme may mainly include:

[0073] (1) A treatment plan was designed to identify factors related to the efficacy of HCC conversion therapy. This treatment plan can be used to deeply analyze the CT image data features of HCC patients before conversion therapy, such as tumor shape, texture, density distribution, liver volume, and grayscale co-occurrence matrix. The chi-square test and standard deviation were used to screen out the imaging features that are strongly correlated with the success of the surgery. That is, the features that are highly correlated with the success rate of the surgery (such as tumor shape, texture, and liver volume) were accurately screened out to reveal the potential relationship between medical imaging data and clinical prediction and prognosis results, providing a key basis for subsequent subclassification and prediction.

[0074] (2) After screening out the features with a high correlation with the surgical success rate, the discreteness of the corresponding imaging features is measured based on the standard deviation, and the box plot is combined with the processing of excluding outliers to select appropriate subclassification indicators to ensure the effectiveness and distinguishing ability of the selected imaging features in subclass classification.

[0075] (3) Based on the above subclassification indicators, unsupervised system clustering is used for the preliminary labeling of unlabeled imaging data (i.e., unlabeled samples), solving the problem of unlabeled samples initially having no subclass labels, and providing a labeled data basis for subsequent supervised learning. After the subclass labeling is completed, the random forest model can be trained to facilitate the subsequent use of the random forest model for efficacy type prediction, thereby achieving a personalized assessment of the patient's surgical success rate, and then exploring the best suitable population for advanced HCC patients to carry out immune combined with targeted drug conversion therapy, optimize the allocation of medical resources, and improve patient survival, providing strong support for promoting the development of liver oncology.

[0076] Among them, the corresponding RF algorithm is an integrated learning based on DT. The RF algorithm first uses the Bootstrap (a front-end framework) method to extract the training set from the original sample set, and then trains the DT model on each training set. Finally, the category or one of the categories with the most votes from all base classifiers is the final category, which has good interpretability.

[0077] This processing can combine unsupervised learning with supervised learning to give full play to the advantages of supervised learning in subclass classification and prediction, improve the accuracy and interpretability of the entire evaluation method, and then provide a corresponding interpretable HCC conversion treatment effect classification evaluation implementation method.

[0078] From the above description, it can be seen that the implementation scheme for evaluating the efficacy of liver cancer conversion therapy based on system clustering and random forest provided in the embodiment of the present invention, its specific implementation process can mainly include: the analysis process of factors related to surgical success rate based on chi-square test, the processing process of significant feature selection based on standard deviation, and the processing process of subclassification of liver cancer conversion therapy effect based on system clustering-random forest.

[0079] For ease of understanding, the specific implementation methods of the above three processing processes will be described in detail below.

[0080] 1. Analysis of factors related to surgical success rate based on chi-square test

[0081] In this process, the chi-square test was used to analyze the differences between the success of the conversion therapy surgery and the characteristics of tumor shape, texture, density distribution, liver volume and gray-level co-occurrence matrix (GLCM) to determine the corresponding factors related to the success rate of the surgery. Figure 1 As shown in the figure, the specific implementation process includes:

[0082] Step 11, collecting, extracting features and preprocessing CT image data;

[0083] In this step, a GE Gemstone Spectrum 256-row CT scanner can be used to perform an upper abdominal CT scan on a liver cancer patient; the corresponding scanning parameters may include: voltage 80-120KV, current 100-300mA, rotation time 0.3-0.5s, detector combination 0.625×64mm, pitch 0.8-0.9, bed entry speed 30-50mm / s, layer thickness 3-7mm, spacing 3-5mm; for example, the corresponding scanning parameters may be selected, but not limited to, as follows: voltage 120KV, current 200mA, rotation time 0.5s, detector combination 0.625×64mm, pitch 0.894, bed entry speed 47.5mm / s, layer thickness 5mm, spacing 5mm;

[0084] The feature extraction process in this step can be implemented by relying on open source radiomics feature extraction platforms and feature extraction software. Specifically, Python+RADIOMICS can be used to extract high-dimensional feature data from the collected CT image data to determine the variables that affect the success rate of the operation. Specifically, the extracted features can include tumor shape and size, texture, density distribution, liver volume size, gray-level co-occurrence matrix (GLCM), local binary pattern, etc.

[0085] The preprocessing in this step includes checking data integrity and deleting missing data to avoid affecting subsequent analysis.

[0086] Reference Figure 1 As shown, the image feature factors are extracted after data preprocessing of the adopted CT image, and then the corresponding image feature is judged whether it is a categorical variable. If so, it is processed to determine the variables affecting the success rate of the operation, and step 12 is executed. Otherwise, the corresponding multi-factor variance analysis is performed to process the continuous variables in the image features; wherein the categorical variable is used to represent a category or classification, and there is no order or quantity relationship between the data. For example, the shape of the tumor is a categorical variable;

[0087] Step 12, determining corresponding significant image features based on correlation analysis of chi-square test;

[0088] The chi-square value is determined by the degree of deviation between the actual observed values ​​and the theoretical inference values ​​of the statistical sample. The larger the chi-square value, the greater the deviation between the two; conversely, the smaller the deviation between the two; if the two values ​​are completely equal, the chi-square value is 0, indicating that the theoretical value is completely consistent.

[0089] Assume that variable X represents whether the operation is successful or not; variable Y represents tumor shape, texture, density distribution and liver volume, etc., then SPSS (social science statistical software) can be used to perform interactive analysis on variables X and Y, and determine the corresponding P value based on the chi-square test, and then judge the significance of the feature based on the P value, and perform corresponding analysis of the correlation of variables, so as to determine the strongly correlated imaging features related to the success rate of the operation, that is, the significant imaging features;

[0090] The following example illustrates the process of determining the strongly correlated imaging features related to the success rate of surgery. It is assumed that the following conclusions can be drawn based on the analysis results:

[0091] The analysis determined that the significance P value of the correlation between surgical success and tumor shape was 0.0035, rejecting the null hypothesis, so there was a significant difference, that is, tumor shape was a strongly correlated imaging feature related to surgical success rate;

[0092] The analysis determined that the significance P value of the correlation between surgical success and texture was 0.020, rejecting the null hypothesis and showing a significant difference, that is, texture is a strongly correlated imaging feature related to surgical success rate;

[0093] The analysis determined that the significance P value of the correlation between surgical success and density distribution was 0.507, and the null hypothesis was accepted, and there was no significant difference, that is, the density distribution was not a strongly correlated imaging feature related to the surgical success rate;

[0094] The analysis determined that the significant P value of the relationship between surgical success and liver volume was 0.019, rejecting the null hypothesis that there was a significant difference, that is, liver volume was a strongly correlated imaging feature related to surgical success rate.

[0095] Through the above analysis process, it can be concluded that the factors related to the success of HCC resection surgery (i.e., factors related to the success rate of surgery) at least include: tumor shape, texture and liver volume. That is, according to the corresponding significant P values, it can be determined that these factors are more correlated with the success of HCC resection surgery, especially the tumor shape with the smallest significant P value, which has the greatest impact on the success rate of surgery among these three factors.

[0096] 2. Selection of image features with significant differences based on standard deviation

[0097] In this processing, the standard deviation can be used as a measurement standard to calculate the degree of dispersion of each imaging feature in the two primary types of successful surgery and unresectable surgery, and then different influencing features are selected as influencing features with significant differences to divide the surgical effect into subcategories. Among them, each imaging feature in the two primary types of successful surgery and unresectable surgery is the strongly correlated imaging feature (i.e., significant imaging feature) determined by the chi-square test.

[0098] Step 21, calculating the standard deviation of each significant image feature determined by the analysis in the above process (I);

[0099] Specifically, using the standard deviation as a benchmark can reflect the degree of dispersion of each image feature. The standard deviation refers to the arithmetic square root of the arithmetic mean (i.e., variance) of the square of the mean deviation. It is the most commonly used measurement basis for the degree of statistical distribution in probability statistics. The corresponding calculation formula is as follows:

[0100] Population standard deviation: ;Sample standard deviation: ;

[0101] Standard error: ; Sample mean: (1)

[0102] in, Represents the overall mean of an indicator. represents the index value of the i-th sample point, n represents the number of sample data, Represents the mean of the sample; the population standard deviation is often used to measure the degree of data dispersion of the entire population, and the sample standard deviation is used to measure the degree of dispersion of sample data and is usually used to estimate the population standard deviation.

[0103] Step 22, using a box plot to capture outliers in significant image features;

[0104] In order to better perform subclassification, in this step, it is necessary to identify and sort out the outliers in the standard deviation data; the corresponding methods for identifying the corresponding outliers may include: in the box plot, sort the data from small to large; the lower quartile Q1 is the value ranked 25%, the upper quartile Q3 is the value ranked 75%, and the interquartile range , based on the sorting result, the abnormal value is captured and determined according to the predetermined interval parameter; assuming that the predetermined interval parameter (i.e., the specified reasonable interval) is: , then the value of the image feature corresponding to the standard deviation outside the interval determined by the interval parameter is the outlier, and the outlier is the image feature with a smaller standard deviation;

[0105] Step 23, based on the above abnormal values, selecting appropriate image features (i.e. image features with significant differences) from the significant image features as indicators for subclassifying surgical effects;

[0106] Considering the rationality of the subclassification model, the quartiles of the standard deviation sequence can be calculated according to the box plot in step 22. The more obvious the difference is, the greater the difference between different classes is, which is helpful for secondary classification. On the contrary, the smaller the difference in standard deviation is (all 0), the more difficult it is to reflect the difference between the data. Therefore, according to the distribution data of the standard deviation, the corresponding selection and determination of the processing process of the image features with significant differences includes:

[0107] The significant image features corresponding to the outliers are selected and removed from the significant image features to exclude outliers with smaller standard deviations, thereby ensuring that the selected features have good distinguishing ability for classification, and the remaining other significant image features are selected as the image features with significant differences to be used as image features for subclass division indicators, that is, the selected image features with significant differences can be used as influencing features for subsequent subclass division.

[0108] 3. Classification of liver cancer conversion therapy effect subcategories and corresponding conversion therapy efficacy evaluation process based on system clustering and random forest

[0109] In the processing, the CT images are divided into various subclass clusters as unlabeled samples by unsupervised system clustering, and a label is defined for each subclass cluster as the label value of each sample included in the unlabeled sample; wherein the subclass includes a subclass of successful surgery and a subclass of non-switchable type; then, each sample included in the unlabeled sample and the label value determined for it are used as a training set, and a random forest model is trained by supervised learning; the random forest model is used to predict and evaluate the efficacy of liver cancer conversion therapy for patients with unknown types of liver cancer;

[0110] Reference Figure 2 As shown, the corresponding process of liver cancer conversion therapy effect subclassification based on system clustering and random forest and the corresponding conversion therapy efficacy evaluation can include:

[0111] Step 31, using unsupervised system clustering to obtain sample labels for unlabeled samples;

[0112] Since the basis for subclassification is not given, each unlabeled sample (i.e., the unlabeled indicator data of CT images) has no subclass label, so it is necessary to first cluster the unlabeled data using system clustering or K-means clustering to perform subclass division and label determination;

[0113] Specifically, the image features with significant differences selected previously can be used as the basis for judging the Euclidean distance in further subclass division, and the optimal K value can be determined by using the elbow method to determine the number of subclass clusters for clustering as K. In this way, the indicator data as samples can be divided into various subclass clusters, and a label is defined for each subclass cluster as the corresponding sample label, such as 1, 2, 3...K subclasses; further, Figure 3 and Figure 4 As shown in the figure, the surgical success type can be clustered into two subtypes, namely high success rate and moderate efficacy; the unresectable HCC type can be clustered into three subtypes, namely, the patient's general condition cannot withstand surgical trauma, the remaining liver volume is insufficient, and the efficacy after resection is poor. The poor efficacy after resection means that the technology can be resected, but the resection cannot achieve a better efficacy than non-surgical treatment. The silhouette coefficient is used to evaluate the quality of clustering.

[0114] Step 32, after the operation of step 31, each unlabeled sample obtains a label value, and then the supervised learning method can be used to train the supervised learning random forest model by taking the corresponding samples and label values ​​as training sets;

[0115] In this step, first, the image features used as the subclass division index can be subjected to multiple random sampling of samples and features, and the area where the subset obtained by each sampling is located can be divided; then, based on the area where the subset is located determined by the division, the optimal feature number and the split point value corresponding to the subset are found by traversing all samples and features; when the number of samples contained in the leaf nodes of the subset is less than a predetermined value, the decision tree DT model of the subset is obtained; finally, when the corresponding multiple DT models are constructed for all subsets obtained by multiple samplings, the random forest model can be constructed based on the multiple DT models;

[0116] Specifically, the training process of the corresponding random forest model may include the following processing steps:

[0117] (321) The bootstrap random sampling method with replacement and the RSM (Random Subspace Method) algorithm of randomly selecting feature subsets were used to analyze the independent variables that affect the subclassification of surgical effects. conduct The samples and features are randomly sampled, and the Secondary generation subset The process is as follows:

[0118] (2)

[0119] in, is a random subspace function, is a random sampling function; Indicates the number of features selected by the sub-training set, usually ; represents the number of samples, Indicates the number of features in the independent variable;

[0120] (322) Division The DT is constructed in the area where it is located, and the optimal variable number (that is, the optimal feature number) and the split point value are found by traversing all samples and features. , the optimization description is as follows:

[0121] (3)

[0122] in, Represents the optimal variable number (which variable is selected as the basis for dividing the data); Represents the split point value (the split point value means that after selecting the optimal variable, a specific threshold is determined to divide the data into two subsets); and Indicated in and The measured value of an image feature of a region, and express and The average of the measurements in the area, is the sample number threshold of the leaf node; st is used to describe the constraints in the optimization problem, that is, and The constraints that need to be met; the measurement value of a certain image feature refers to the specific value of a certain feature in a certain area. For example, if the image feature is the liver volume, the measurement value is the specific size of the liver volume; if the image feature is the tumor shape, the measurement value is the specific shape of the tumor;

[0123] Use the following criteria to divide Area and determine the corresponding output value:

[0124] (4)

[0125] in, Represents the optimal variable number (which variable is selected as the basis for dividing the data); Represents the split point value (the split point value means that after selecting the optimal variable, a specific threshold is determined to divide the data into two subsets); and Indicated in and The measurement of an image feature for an area.

[0126] (323) Yes and Repeat the above step (322) until the number of samples contained in the leaf node is less than the predetermined value, that is, the number of samples in the leaf node is less than the set threshold. ; Accordingly, the input space is divided into The obtained DT model is recorded as follows:

[0127]

[0128] (5)

[0129] in, represents the decision tree model on the jth subset, Represents the number of divided areas; The area contains samples indivual; express In the region In the subset a truth value; represents an indicator function, when When present ,otherwise ; represents the number of features selected by the sub-training set, that is, the number of features selected by the sub-training set at the jth split. Indicates the candidate split point value under a certain feature.

[0130] (324) Repeat the above steps (321) to (323) Next, the RF sub-model is obtained based on the following formula:

[0131] (6)

[0132] in, Represents the integrated random forest model; Represents the decision tree model obtained on the jth subset.

[0133] Based on the RF sub-model determined by the above formula, the trained random forest model is obtained;

[0134] Step 33, using the random forest model obtained after training to predict the efficacy evaluation;

[0135] Specifically, the influencing feature data (CT images) of patients with unknown types of liver cancer can be input into the trained random forest model, and after corresponding processing, the prediction results of the efficacy evaluation of liver cancer conversion therapy can be obtained.

[0136] In summary, the HCC conversion treatment effect classification evaluation implementation scheme based on system clustering and random forest provided in the embodiment of the present invention adopts the analysis and determination of the implementation scheme of the factors related to the prognosis effect of HCC conversion treatment based on the chi-square test, so that the correlation between the surgical success rate and the characteristics of tumor shape, texture, density distribution, liver size, gray level co-occurrence matrix (GLCM) and local binary pattern can be analyzed and determined, providing guidance for subsequent decision-making. The embodiment of the present invention also uses the standard deviation as a measurement standard to reflect the degree of discreteness of each image feature, so as to help select appropriate features as indicators for subclass division; further, the outliers in the image features can be identified through the box plot to exclude outliers with smaller standard deviations, so that the subclass classification is more reasonable.

[0137] Furthermore, the embodiment of the present invention also combines unsupervised learning with supervised learning, that is, first use the unsupervised system to cluster the unlabeled samples for labeling, and then use supervised learning to train the corresponding random forest model to improve the accuracy and interpretability of the efficacy evaluation prediction.

[0138] In practical applications, the embodiment of the present invention selects a batch of high-quality submillimeter CT data of HCC patients, and uses the silhouette coefficient to evaluate the quality of clustering. The application results show that an effect of 0.8893 can be achieved. Furthermore, a comparative experiment of the classification evaluation of the efficacy of conversion therapy is performed using the Kmeans-CART model and the technical solution provided by the embodiment of the present invention. The comparative experimental results show that the technical solution provided by the embodiment of the present invention achieves an accuracy of 89.32%, which is higher than the 76% accuracy of K-means-CART, thereby confirming the effectiveness of the technical solution provided by the embodiment of the present invention. Moreover, the technical solution provided by the embodiment of the present invention has better interpretability than a single clustering method, and has certain guiding significance for exploring the classification rules of the efficacy of conversion therapy for liver cancer. In addition, during the experiment, a perturbation experiment with an interval of (=1%, 10%, ..., 30%) is set for each image feature. The corresponding experimental results show that the accuracy of the model changes slightly before and after the interference. When the disturbance increases to 30%, the accuracy drops to 86%, thereby proving that the technical solution provided by the embodiment of the present invention has good robustness.

[0139] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with the technical field within the technical scope disclosed in the present invention should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims. The information disclosed in the background technology section of this article is only intended to deepen the understanding of the overall background technology of the present invention, and should not be regarded as an admission or in any form that the information constitutes prior art known to those skilled in the art.

Claims

1. A method for evaluating the efficacy of liver cancer conversion therapy, characterized in that: include: In the computer tomography CT images of liver cancer patients, image features with significant differences pre-selected based on significant image features are used as the basis for judging the Euclidean distance in subclass division, and the CT images are divided into various subclass clusters as unlabeled samples by unsupervised system clustering; wherein, the method for determining the significant image features includes: collecting and extracting image features corresponding to surgical success rate-related factors in the CT images of liver cancer patients, and using the chi-square test analysis method to statistically calculate the degree of deviation between the actual observed value and the theoretical inferred value of the image features as samples, and then determining the significant image features according to the statistical results; and the discrete degree of each image feature is measured by the standard deviation and the image features with significant differences are selected; the image features include tumor shape, texture, density distribution, liver volume and grayscale co-occurrence matrix; A label is defined for each subclass cluster as the label value of each sample contained in the unlabeled sample; wherein the subclass includes a subclass of successful surgery and a subclass of unresectable surgery; wherein the subclass of successful surgery includes two subclasses: high success rate and moderate efficacy; the subclass of unresectable surgery includes three subclasses: inability to withstand surgical trauma, insufficient residual liver volume, and poor efficacy after resection; Each sample included in the unlabeled sample and the label value determined for it are used as a training set, and a random forest model is trained by supervised learning; the random forest model is used to predict and evaluate the efficacy of liver cancer conversion therapy for patients with unknown types of liver cancer.

2. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 1, characterized in that: The method for obtaining the image features with significant differences pre-selected based on the significant image features includes: Among the obtained imaging features of liver cancer patients, the corresponding significant imaging features are determined based on the factors related to the surgical success rate, and the standard deviation of each significant imaging feature is calculated; Based on the standard deviation, using a box plot to capture outliers in the significant image features; Based on the captured outliers in the significant image features, the image features with significant differences are selected and determined as image features for subclass division indicators.

3. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 1, characterized in that: The image features include: tumor shape, size, texture and density distribution, as well as liver volume, gray level co-occurrence matrix GLCM and local binary pattern.

4. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 2, characterized in that: The scanning parameters of the CT image include: Voltage 80-120KV, current 100-300mA, rotation time 0.3-0.5s, detector combination 0.625×64mm, pitch 0.8-0.9, bed feed speed 30-50mm / s, layer thickness 3-7mm, spacing 3-5mm.

5. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 1, characterized in that: The processing of determining the significant image features according to the statistical results includes: The imaging feature and the success or failure of the operation were taken as two variables, and an interactive analysis was performed on the two variables based on the social science statistical software SPSS. According to the significance P value of each imaging feature obtained by the interactive analysis, it was determined whether the corresponding imaging feature was the significant imaging feature.

6. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 2, characterized in that: The standard deviation includes: overall standard deviation, sample standard deviation, standard error and sample mean, and the corresponding calculation formula is: Population standard deviation for: ; Sample standard deviation for: ; Standard error for: ; Sample mean for: ; in, Represents the overall mean of an indicator. represents the index value of the i-th sample point, n represents the number of sample data, Represents the mean of the sample; the population standard deviation is often used to measure the degree of data dispersion of the entire population, and the sample standard deviation is used to measure the degree of dispersion of sample data.

7. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 2, characterized in that: The process of using a box plot to capture outliers in the significant image features includes: In the box plot, the standard deviations are sorted from small to large; the lower quartile Q1 is the value ranked 25%, the upper quartile Q3 is the value ranked 75%, and the interquartile range ; Based on the box plot, capturing and determining the abnormal value according to a predetermined interval parameter; The interval parameters include: ; The process of capturing and determining the abnormal value includes: identifying the image feature corresponding to the standard deviation outside the interval corresponding to the interval parameter as an abnormal value; The process of selecting and determining the image features with significant differences includes: The significant image features corresponding to the outliers are selected and removed from the significant image features, and the remaining other significant image features are selected as the image features with significant differences to be used as image features for subclass division indicators.

8. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to any one of claims 1 to 7, characterized in that: The process of training the random forest model using supervised learning includes: Performing multiple random sampling of samples and features on the image features used as subclass division indicators, and dividing the area where the subsets obtained by each sampling are located; Based on the area where the subset is located, the optimal feature number and split point value corresponding to the subset are found by traversing all samples and features; when the number of samples contained in the leaf nodes of the subset is less than a predetermined value, a decision tree DT model of the subset is obtained; After the corresponding multiple DT models are constructed for all subsets obtained through multiple samplings, the random forest model is constructed based on the multiple DT models.

9. The method for evaluating the therapeutic efficacy of liver cancer conversion therapy according to claim 8, characterized in that: The process of obtaining the decision tree DT model of the subset includes: The independent variables corresponding to the image features of the subclass division index are analyzed by using the bootstrap random sampling method with replacement and the RSM algorithm for randomly selecting feature subsets. conduct The samples and features are randomly sampled. times is multiple times, the first Secondary generation subset The process is as follows: ; in, is a random subspace function, is a random sampling function; represents the number of features selected in the sub-training set, m represents the sequence number of the feature in the independent variable, , where M is the number of features in the independent variable and N is the number of samples; Divide the subset The area where it is located, and find the optimal feature number and split point value by traversing all its samples and features , the value for: ; in, and Indicated in and The measured value of an image feature of a region, and express and The average of the measurements in the area, is the sample number threshold contained in the leaf node, st refers to and The constraints that need to be met; represents the optimal variable number; represents the value of the segmentation point; the measurement value of a certain image feature refers to the specific value of a certain image feature in a certain area; The sub-set The area and the corresponding output value are determined using the following criteria: ; For the area and Repeat the steps of finding the optimal feature number and split point value until the number of samples in the leaf node of the subset is less than the set threshold ; Accordingly, the input space is divided into The DT model of this subset is obtained as follows: ; ; in, represents the decision tree model on the jth subset, Represents the number of divided areas; The area contains samples indivual; express In the region In the subset a truth value; represents an indicator function, when When present ,otherwise ; represents the number of features selected by the sub-training set, Indicates the candidate split point value under a certain feature; And the process of constructing the random forest model based on multiple DT models includes: exist Repeat the above process to obtain the decision tree DT model of this subset times, obtained After the DT model, the corresponding random forest sub-model is: ; in, Represents the integrated random forest model.

Citation Information

Patent Citations

  • MSWI process dioxin emission soft measurement method based on missing data filling

    CN114970353A

  • Prediction method for rectal cancer chemoradiotherapy reaction effect based on feature selection method

    CN117038066A

  • Construction method of minimally invasive surgery risk assessment model based on craniocerebral tumor

    CN119361081A