Thyroid lymphoma prediction method and device based on machine learning

By segmenting and extracting lesion areas on thyroid ultrasound images, and using random forest models for feature screening and training, the accuracy of early diagnosis of thyroid lymphoma was solved, and efficient disease risk prediction and early screening were achieved.

CN120298412AActive Publication Date: 2025-07-11PEKING UNIVERSITY THIRD HOSPITAL (THE THIRD CLINICAL MEDICAL SCHOOL OF PEKING UNIVERSITY)
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510787635.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-07-11
Estimated Expiration
2045-06-13

AI Technical Summary

Technical Problem

In the prior art, it is difficult to accurately distinguish between benign nodules and malignant lymphomas in the early diagnosis of thyroid lymphoma, resulting in misdiagnosis or misdiagnosis. In addition, deep learning technology lacks systematic research and effective solutions in disease risk prediction.

Method used

By segmenting and labeling the lesion area of the thyroid ultrasound image, tumor imaging features were extracted, combined with zero importance analysis and correlation screening, the thyroid lymphoma prediction model was trained using a random forest model to perform feature dimensionality reduction and redundancy, and improve the accuracy and efficiency of the model.

Benefits of technology

It realizes accurate prediction of thyroid lymphoma, improves the accuracy and efficiency of diagnosis, and provides clinicians with support tools for early warning and precise prevention and control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298412A_ABST
    Figure CN120298412A_ABST
Patent Text Reader

Abstract

The invention relates to a thyroid lymphoma prediction method and device based on machine learning, belongs to the technical field of image processing, and solves the problem of lack of a thyroid lymphoma prediction method based on deep learning in the prior art. The method comprises the following specific steps: marking a lesion type label containing lymphoma for a thyroid ultrasound image in which a lesion area is segmented to obtain an original data set; extracting tumor iconography features of the original data set to obtain initial features, and performing zero importance analysis and correlation screening in sequence to obtain final screening features with high importance and low correlation; based on the final screening features and the marked lesion type labels, a thyroid lymphoma prediction model is obtained through machine learning training; feature extraction is carried out on the to-be-detected thyroid ultrasound segmentation image with the suspected lesion area segmented, and corresponding final screening features are obtained; and based on the corresponding final screening characteristics, realizing accurate prediction of thyroid lymphoma by using a thyroid lymphoma prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to a method and device for predicting thyroid lymphoma based on machine learning. Background Art

[0002] Thyroid lymphoma is a malignant tumor that poses a significant threat to human health. This disease not only has a high incidence and mortality rate but also causes a heavy economic and psychological burden on patients and their families due to the high cost and complex process of treatment. Traditional treatment methods mainly include chemotherapy, radiotherapy, and surgical resection. However, due to the non-obvious early symptoms of thyroid lymphoma, patients are often diagnosed at an advanced stage of the disease. At this time, the tumor may have widely invaded surrounding tissues or metastasized distantly, resulting in poor treatment effects, poor prognosis, and low survival rate for patients. Therefore, how to accurately predict the risk of thyroid lymphoma at an early stage and achieve early detection, early diagnosis, and early treatment is the key to improving the cure rate and prognosis of patients. In the field of medical imaging, ultrasound examination is the preferred method for screening and diagnosing thyroid diseases, with advantages such as non-invasive, convenient, and highly repeatable. However, for the diagnosis of thyroid lymphoma, relying solely on the morphological features of ultrasound images has certain limitations, making it difficult to accurately distinguish between benign nodules and malignant lymphomas and prone to misdiagnosis or missed diagnosis. In recent years, with the rapid development of deep learning technology, its application in medical image analysis has become increasingly widespread, providing new ideas and methods for the early diagnosis and disease prediction of thyroid lymphoma.

[0003] In the clinical research and practice of thyroid lymphoma, the precise segmentation and accurate diagnosis of lesions are the crucial first step. Currently, researchers have made significant progress, especially in using deep learning technology to assist in the analysis of ultrasound images, which can more accurately identify and delineate thyroid lymphoma and play an important role in improving the accuracy and efficiency of diagnosis. However, there are relatively few deep learning studies on thyroid lymphoma at present, especially in using deep learning technology for disease prediction. It is still in its infancy, and there is a lack of systematic research and effective solutions for using deep learning models to predict individual disease risks and achieve early screening and identification of high-risk populations. Summary of the Invention

[0004] In view of the above analysis, the embodiments of the present invention aim to provide a method and device for predicting thyroid lymphoma based on machine learning to solve the problem of the blank in thyroid lymphoma prediction based on deep learning in the prior art.

[0005] The object of the present invention is mainly achieved through the following technical solutions: On the one hand, the embodiments of the present invention provide a method for predicting thyroid lymphoma based on machine learning, including the following steps: For the thyroid ultrasound images that have segmented the lesion area, label the lesion type tags including lymphoma to obtain the original dataset; Extract the tumor imaging features of the original dataset to obtain the initial features, and perform zero importance analysis and correlation screening in sequence to obtain the final screening features with high importance and low correlation; Based on the final screening features and the labeled lesion type tags, use machine learning to train a thyroid lymphoma prediction model; Extract features from the thyroid ultrasound segmentation image to be measured that has segmented the suspected lesion area to obtain the corresponding final screening features; Based on the corresponding final screening features, use the thyroid lymphoma prediction model to predict thyroid lymphoma.

[0006] Furthermore, the zero importance analysis of the initial features includes: Train a random forest model based on the original dataset, and calculate the importance score of each initial feature under the condition of the true label; wherein, the true label is the lesion type label corresponding to the lesion area in the thyroid ultrasound segmentation image; Randomly scramble the lesion type tags to obtain a pseudo-label dataset, and retrain the random forest model based on the pseudo-label dataset, and calculate the importance score of each initial feature under the condition of the pseudo-label; Compare the distributions of the importance scores of each initial feature under the conditions of the true label and the pseudo-label to obtain the zero importance score of this initial feature; Screen out the features with zero importance scores greater than the preset zeroing threshold among all the initial features to obtain the high-importance features.

[0007] Furthermore, the correlation screening of the high-importance features to obtain the final screening features includes: Calculate the correlation between each high-importance feature to obtain a correlation matrix; Sort the high-importance features based on the zero importance scores; Traverse each high-importance feature in the sorted order and combine with the correlation matrix for pairwise comparison, and screen out the features with correlation lower than the preset correlation threshold to obtain the final screening features.

[0008] Furthermore, use the Gini index to calculate the importance score of each initial feature, and obtain the zero importance score of the initial feature based on the following formula: , wherein, is the zero importance score of the initial feature i; is the importance score of this initial feature under the condition of the true label; is the importance score of the initial feature under the pseudo-label condition; P() represents the probability operation.

[0009] Furthermore, the main categories of the initial features include first-order features, shape features, and high-level texture features. Based on the final screening features and the labeled lesion type labels, a thyroid lymphoma prediction model is trained using a random forest.

[0010] Furthermore, an automatic sampling technique is adopted to train a thyroid lymphoma prediction model using a random forest, including: Based on the final screening features and the labeled lesion type labels, a training data set is obtained, and multiple sub-data sets are generated by sampling the training data set with replacement; Based on each sub-data set, the corresponding decision trees in the thyroid lymphoma prediction model based on the random forest are independently trained; Based on the following formula, the prediction results of each decision tree are integrated to obtain the final prediction result of thyroid lymphoma: , where, is the final prediction result of thyroid lymphoma; is the prediction result of the i-th decision tree; is the total number of decision trees.

[0011] Furthermore, when independently training the corresponding decision trees, the sample features and labels in the sub-data sets are sampled with replacement, and the corresponding training subsets are generated based on the following formula: , where, is the training subset corresponding to the i-th decision tree; respectively represent the training samples independently sampled from the sample feature distribution and the label distribution .

[0012] Furthermore, the tumor imaging features of the original data set are extracted, including: Using the gray histogram, the global intensity features of the thyroid ultrasound segmentation images in the original data set are extracted to obtain first-order features; Using the area and perimeter of the segmented region, geometric shape features are extracted to obtain shape features; Using the gray-level co-occurrence matrix, gray-level dependence matrix, and gray-level run length matrix, local structure information is extracted to obtain high-level texture features.

[0013] On the other hand, an embodiment of the present invention provides a thyroid lymphoma prediction device based on machine learning, including: An image acquisition module, configured to collect ultrasound images including the thyroid region; The segmentation and annotation module is used to segment the lesion area on the thyroid ultrasound image to generate a thyroid ultrasound segmentation image; it is also used to annotate the lesion type label containing lymphoma on the thyroid ultrasound segmentation image; The feature extraction module is used to extract the tumor imaging features of the thyroid ultrasound segmentation image to obtain initial features, and perform zero importance analysis and correlation screening on the initial features in sequence to obtain final screening features with high importance and low correlation; it is also used to extract features from the thyroid ultrasound segmentation image to be measured with suspected lesion areas segmented to obtain corresponding final screening features; The prediction module is used to train a thyroid lymphoma prediction model based on the final screening features and the annotated lesion type labels using machine learning; it is also used to predict thyroid lymphoma based on the corresponding final screening features using the thyroid lymphoma prediction model; The human-computer interaction module is used to output the prediction result for the loaded imaging and feature data using visualization software and interface.

[0014] Furthermore, performing zero importance analysis on the initial features using the feature extraction module includes: Obtaining an original data set based on the thyroid ultrasound segmentation map annotated with lesion type labels, training a random forest model based on the original data set, and calculating the importance score of each initial feature under the condition of the true label; wherein, the true label is the lesion type label corresponding to the lesion area in the thyroid ultrasound segmentation image; Randomly scrambling the lesion type labels to obtain a pseudo-label data set, and retraining the random forest model based on the pseudo-label data set, and calculating the importance score of each initial feature under the condition of the pseudo label; Comparing the distribution of the importance scores of each initial feature under the conditions of the true label and the pseudo label to obtain the zero importance score of this initial feature; Screening out the features with zero importance scores greater than a preset zeroing threshold among all the initial features to obtain high importance features.

[0015] Compared with the prior art, the present invention can at least achieve one of the following beneficial effects: 1. The present invention proposes to segment and annotate the lesion area based on the thyroid ultrasound image, perform feature extraction through the extracted tumor imaging category features, combine zero importance analysis and correlation analysis, and use the highly representative and general tumor imaging features extracted as input, and combine machine learning training to obtain a thyroid lymphoma prediction model; extract corresponding final screening features from the thyroid ultrasound segmentation map to be measured, and use the thyroid lymphoma prediction model to achieve accurate prediction of thyroid lymphoma.

[0016] 2. To make up for the limitations of single features, first-order features, shape features, and high-level texture features are extracted based on thyroid ultrasound segmentation images to improve the generalization ability of the model. Using the zero importance analysis and correlation judgment of the random forest model, the extracted multiple features are further dimension-reduced and redundant features are removed, and features with high correlation with the label are screened out, improving the accuracy and efficiency of model prediction.

[0017] 3. The performance of prediction models constructed by the multi-layer perceptron algorithm and the random forest is compared, and a prediction model based on the random forest is selected to achieve the prediction of thyroid lymphoma, improving the interpretability of the model. Combining the processing of redundant input features, the computational cost brought by the random forest is reduced, and the prediction speed of the model is improved, providing a powerful decision support tool for clinicians to achieve early warning and precise prevention and control of thyroid lymphoma.

[0018] In the present invention, the above technical solutions can also be combined with each other to achieve more preferred combination schemes. Other features and advantages of the present invention will be described in the subsequent specification, and some advantages can be made obvious from the specification or understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the content specifically pointed out in the specification and the drawings. Description of the Drawings

[0019] The drawings are only for the purpose of showing specific embodiments and are not considered as a limitation of the present invention. Throughout the drawings, the same reference signs denote the same components; Figure 1 It is a flowchart of the method for predicting thyroid lymphoma based on machine learning according to an embodiment of the present invention; Figure 2 It is a heat map of high-importance and low-correlation features obtained in the feature extraction stage according to an embodiment of the present invention; Figure 3 It is a schematic diagram of the principle of the random forest algorithm according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the principle of the multi-layer perceptron algorithm according to an embodiment of the present invention; Figure 5 It is a confusion matrix result diagram obtained when the random forest prediction model according to an embodiment of the present invention is processed; Figure 6 It is a confusion matrix result diagram obtained when the multi-layer perceptron prediction model according to an embodiment of the present invention is processed; Figure 7 It is an ROC curve diagram obtained when the random forest prediction model according to an embodiment of the present invention is processed; Figure 8 It is an ROC curve diagram obtained when the multi-layer perceptron prediction model according to an embodiment of the present invention is processed. Detailed Embodiments

[0020] The preferred embodiments of the present invention will be specifically described below with reference to the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0021] Embodiment 1 A specific embodiment of the present invention discloses a method for predicting thyroid lymphoma based on machine learning, as Figure 1 shown, which includes the following steps: Step S1: For the thyroid ultrasound images with the lesion areas segmented, label the lesion type labels including lymphoma to obtain the original data set; Step S2: Extract the tumor imaging features of the original data set to obtain the initial features, and perform zero importance analysis and correlation screening in sequence to obtain the final screening features with high importance and low correlation; Step S3: Based on the final screening features and the labeled lesion type labels, use machine learning to train a thyroid lymphoma prediction model; Step S4: Extract features from the thyroid ultrasound segmentation image to be measured with the suspected lesion area segmented to obtain the corresponding final screening features; Step S5: Based on the corresponding final screening features, use the thyroid lymphoma prediction model to predict thyroid lymphoma.

[0022] Through the above method, by using the thyroid ultrasound segmentation images based on the labeled lesion type labels, extracting tumor imaging features, performing zero importance analysis and correlation analysis on the extracted various features, training a thyroid lymphoma prediction model based on the features after removing redundancy, and using the model to predict the thyroid ultrasound segmentation image to be measured, high-precision and high-robustness prediction of thyroid lymphoma is achieved.

[0023] Specifically, in step S1, for the obtained several ultrasound images including the thyroid lymphoma area, divide the lesion area and label the lesion type labels. The specific steps include: S11: Screen the thyroid ultrasound images that meet the requirements: Select thyroid ultrasound images with higher clarity, resolution meeting the requirements and meeting a certain quantity condition; S12: Determine the preliminary range of the lesion area including lymphoma in the thyroid ultrasound image: Combine the thyroid ultrasound image with the diagnosis result of a professional doctor to clarify the approximate range of the lymphoma area and verify its accuracy with the professional doctor; S13: Segment the areas of lesion types such as lymphoma in the thyroid ultrasound image to generate the corresponding thyroid ultrasound segmentation image; Exemplarily, based on real medical clinical data, the basic information description of the cases and the corresponding thyroid ultrasound images are obtained, including 32 cases of primary thyroid malignant lymphoma (PTML), and in addition, there are 4 cases of papillary thyroid carcinoma, 8 cases of follicular thyroid tumors, and 8 cases of other thyroid diseases in the control group, with a total of 52 cases. The obtained thyroid ultrasound images are 145 groups in total. Based on the medical image processing software 3D-Slicer, the lesion area is drawn on the thyroid ultrasound image to obtain the thyroid ultrasound segmentation image with the lesion area segmented.

[0024] S14. Manually label the lesion type tags for the thyroid ultrasound segmentation image with the lesion area segmented, save it and adjust it after review by professional physicians, and finally obtain the labeled segmentation image. Among them, the labeled thyroid ultrasound segmentation images constitute the original data set.

[0025] Specifically, in step S2, before feature extraction, preprocessing such as scaling, normalization, and denoising is performed on the thyroid ultrasound segmentation images in the original data set. In order to be able to capture more comprehensively the complex information related to thyroid lymphoma, based on the preprocessed images, the steps for extracting tumor imaging features and obtaining the final screening features are as follows: S21. Extract various types of tumor imaging features from the original data set to obtain initial features, including the main category features of first-order features, shape features, and high-level texture features; among them, each main category contains multiple sub-category features, that is, the initial features contain multiple sub-category tumor imaging features.

[0026] Exemplarily, the feature extraction is carried out with the help of the third-party pyradiomics package library of python. The specific process of obtaining the tumor imaging features of the thyroid ultrasound segmentation image by enabling various image types and optional features is as follows: S211. According to the segmentation requirements of thyroid lymphoma and the characteristics of thyroid ultrasound images, set the feature parameters for extracting tumor imaging features, including the width of the histogram division interval bin, the standard deviation sigma of the Gaussian filter, the interpolator type, the pixel spacing, and the voxel array offset; Among them, by adjusting the interpolator type, the segmentation image is scaled or rotated to ensure the smoothness of the image and the retention of details; by adjusting the standard deviation sigma of the Gaussian filter, the Gaussian filter is used to smooth the image, remove noise, and at the same time maintain the main structural features of the image to achieve the preprocessing of the image. When extracting features, by adjusting the width parameter of the histogram, the gray level accuracy of the image is set for the texture analysis of the image; adjusting the pixel spacing controls the spatial resolution of the image for the analysis of shape features and texture features.

[0027] S212. Using features such as the gray histogram features, mean, and standard deviation, extract the global intensity features of the thyroid ultrasound segmentation images in the original dataset to obtain first-order features; using the area, perimeter, shape factor, etc. of the segmented regions, extract the geometric shapes of the lesion tumor regions to obtain shape features; using the gray-level co-occurrence matrix (GLCM), gray-level dependence matrix (GLDM), and gray-level run length matrix (GLRLM) to extract the local structure information of the images to obtain high-level texture features. In the above way, enable all sub-categories corresponding to the first-order features, shape features, and high-level texture features in the tumor imaging features, and extract 1042 tumor imaging features. By comprehensively considering various features and using their joint effects, the diagnostic accuracy can be significantly improved, and the deficiency that a single feature cannot provide enough information to ensure the accurate diagnosis of diseases can be solved.

[0028] S22. Use the principal component analysis (PCA) technique to reduce the dimensionality of the feature space and select the feature subset that makes the greatest contribution to the classification task, thereby reducing the computational complexity and preventing overfitting.

[0029] Specifically, based on the initial features of the extracted thyroid ultrasound segmentation images, perform zero importance analysis to obtain high-importance features. The specific steps are as follows: S221. Train a random forest model based on the original dataset, calculate the feature importance scores of each initial feature under the condition of the true label, and record the feature importance distribution. Among them, the true label is the lesion type label corresponding to the lesion region in the thyroid ultrasound segmentation image; after arranging all the initial features according to the importance scores, the feature importance distribution is obtained.

[0030] Exemplarily, use the Gini index to calculate the importance scores of each initial feature. By measuring the purity of the initial features through the Gini index, all initial features can be best segmented. Its calculation formula is: , where, is the importance score of the initial feature i; is the sample probability corresponding to the initial feature i; is the total number of initial feature categories.

[0031] The lower the Gini index, the lower the feature importance score under the condition of the true label, indicating that the feature is purer; the higher the Gini index, the more mixed the feature category distribution. By comparing the effects of different feature Gini indices, the decision tree model can automatically select the most useful features for data partitioning, thereby optimizing the classification results.

[0032] S222. Perform a random permutation operation on the lesion type labels to disrupt the true mapping relationship between features and labels, construct a pseudo-label dataset, and ensure that the pseudo-label data does not contain valid feature information; S223. Retrain the random forest model 100 times based on the pseudo-label dataset, calculate the importance scores of each initial feature under the pseudo-label condition, and record the feature importance distribution under the pseudo-label; S224. Compare the distributions of the importance scores of each feature under the true label condition and the pseudo-label condition, and calculate the zero-importance score of the initial feature; Exemplarily, for each feature , the calculation formula for the zero-importance score is: , where is the zero-importance score of the initial feature i, that is, the cumulative probability of the actual importance value in the random distribution; is the importance score of the initial feature under the true label condition; is the importance score of the initial feature under the pseudo-label condition; P() is the probability operation. The higher the score, the greater the actual contribution of the feature in the classification task.

[0033] It should be noted that usually, a numerical simulation method is used to calculate the probability with a statistic to obtain the zero-importance score. Extract a large number of samples (such as N) from the pseudo-label results of a certain feature, and count the number n of samples with importance scores less than the corresponding "Actual Importance" among these samples, then score = n / N. For example, extract 10,000 samples of the null hypothesis distribution, and 8,000 of them are less than the actual importance value, then score = 8,000 / 10,000 = 0.8.

[0034] S225. Screen out the features with zero-importance scores greater than the preset zeroing threshold among all initial features to obtain high-importance features.

[0035] Exemplarily, let the zeroing threshold be 0.6. According to the zero-importance analysis results, screen out the features with significant importance and optimize the feature set.

[0036] S23. Use correlation screening to select features highly correlated with the lesion type label and remove redundant features among the high-importance features, specifically including: S231. Calculate the correlation between each high-importance feature using the Pearson correlation coefficient to obtain a correlation matrix; S232. Sort the high-importance features based on the zero-importance scores; S233. Traverse each high-importance feature in the sorting order, and combine with the correlation matrix to perform pairwise comparison, filter out the features with the correlation coefficient lower than the preset correlation threshold, and obtain the final screening features with low correlation and high zero-importance scores.

[0037] Exemplarily, arrange the high-importance features in descending order according to the corresponding zero-importance scores to form a zero-importance list, and select each feature in the zero-importance list from high to low for pairwise comparison with other high-importance features in the list. First, take the first feature in the zero-importance list as the selected feature, and judge the correlation coefficient between the second feature in the zero-importance list and the selected feature. If the correlation is higher than the threshold of 0.6, it means that the second feature is highly correlated with the selected feature, and the second feature is deleted from the zero-importance list; otherwise, the second feature is added to the retained feature list; then, judge the correlation between the first feature and the third feature in the zero-importance list; similarly, after judging the correlation between the first feature and all other features in the zero-importance list in turn, then take the second feature retained in the zero-importance list as the selected feature, and judge the correlation between the currently selected feature and the remaining other features in turn. After zero-importance analysis and correlation analysis, 17 features with low correlation and high importance are obtained in the retained feature list, and its heat map is as Figure 2 shown, and finally a more effective feature set is obtained, that is, the final screening features including three categories: first-order features, shape features and high-level texture features.

[0038] Through the above method, feature selection and optimization are carried out on the preliminary features, the most informative tumor imaging features are selected, and the features most relevant to the diagnosis are retained to optimize the performance of the model and reduce the computational complexity.

[0039] Specifically, in step S3, based on the final screening features and the labeled lesion type labels, a thyroid lymphoma prediction model is trained using a random forest.

[0040] Exemplarily, the network structure of the thyroid lymphoma prediction model trained based on the random forest algorithm is as Figure 3 shown. By introducing the random forest algorithm, an efficient prediction model for the prevalence of thyroid lymphoma is constructed. This model integrates multiple decision trees and integrates multi-dimensional features in the data, so as to focus on the core factors of lymphoma prevalence, specifically including: S31. Build the basic structure of the random forest model and set the decision tree parameters; S32. Use the Bootstrap Sampling (Bagging) module to reduce the variance of the random forest model and avoid the overfitting problem of a single decision tree. The specific steps are as follows: S321. Obtain a training data set based on the final screening features and the labeled lesion type tags, and generate multiple sub-data sets by sampling the training data set with replacement. Each sub-data set is obtained in the following manner: , where is the i-th sub-data set; represents the number of samples in the i-th sub-data set; represent the features and labels of the samples respectively.

[0041] It should be noted that the thyroid lymphoma prediction model is a binary classification model. Adjust the labels of the final screening features input, mark the final screening features of patients with thyroid lymphoma as 1, and the rest as 0.

[0042] Exemplarily, adjust the label of the ultrasound image of a patient with thyroid lymphoma to 1, that is, adjust the label corresponding to the obtained final screening features to 1; adjust the label of the ultrasound image of other thyroid diseases in the control group to 0, that is, adjust the label corresponding to the obtained final screening features to 0; finally, 76 are marked with 1 and 69 are marked with 0. These 145 groups of thyroid ultrasound segmentation images are split into an 80% training data set and a 20% validation set.

[0043] S322. Based on each sub-data set, independently train the corresponding decision tree in the thyroid lymphoma prediction model based on random forest; Exemplarily, set the training process of each decision tree as: , where represents the i-th decision tree, which is trained through the sub-data set .

[0044] Furthermore, during the training process of each tree, adopt a sampling method with replacement for the sample features and labels in the sub-data set. The generation process of the training subset is expressed as:

[0045] where is the training subset corresponding to the i-th decision tree; respectively represent the training samples independently sampled from the sample feature distribution and the label distribution .

[0046] S323. Integrate the prediction results of each decision tree based on the following formula to obtain the final prediction result of thyroid lymphoma: , where To predict the final outcome of thyroid lymphoma; is the prediction result of the i-th decision tree; is the total number of decision trees.

[0047] In order to compare the effects of different machine learning models on the prediction of thyroid lymphoma, as another optional scheme, a thyroid lymphoma prediction model is obtained by training a multi-layer perceptron based on the final screening features and the annotated lesion type labels.

[0048] For example, 80% of the 145 re-labeled data sets are used as training sets and 20% as test sets. The network structure of the thyroid lymphoma prediction model based on the multi-layer perceptron (MLP) algorithm is as follows: Figure 4 As shown in Figure 2, the model uses a multi-layer fully connected network to perform deep learning on features and integrate nonlinear relationships at different levels.

[0049] When training the above two prediction models, confusion matrices are introduced to evaluate model performance, and training parameters are continuously adjusted to make the training results stable. The confusion matrix results are as follows: Figure 5 and Figure 6 shown.

[0050] The receiver operating characteristic (ROC) curve was used to evaluate the two thyroid lymphoma prediction models obtained by machine learning. Figure 7 , Figure 8 As shown in the figure, the accuracy of the random forest model is 0.90, the average accuracy of the five-fold cross validation is 0.84, the average accuracy of the ten-fold cross validation is 0.81, and the area under the ROC curve and the coordinate axis AUC value reaches 0.90. The accuracy of the multi-layer perceptron model is 0.86, and the area under the ROC curve and the coordinate axis AUC value reaches 0.90.

[0051] The prediction results are clearly shown in Table 1. The two prediction models show good prediction performance. However, the random forest model can effectively process high-dimensional data and avoid overfitting. It is more interpretable than the multi-layer perceptron and can be used as a preferred solution. The random feature selection of the Bagging module reduces the correlation between trees, ensuring high accuracy and robustness of the prediction of thyroid lymphoma. In addition, the pre-redundancy removal of the input multi-dimensional features further improves the efficiency of the model, providing strong support for the formulation of personalized treatment strategies.

[0052] Table 1

[0053] Specifically, in step S4, the suspected lesion area is segmented in the thyroid ultrasound image to be measured to obtain the corresponding thyroid ultrasound segmentation image, and then the corresponding features are extracted according to the category of the final screening features to obtain the corresponding final screening features.

[0054] Specifically, in step S5, the obtained corresponding final screening features are input into the thyroid lymphoma prediction model to obtain the prediction result of the thyroid lymphoma disease condition; the prediction result is a binary output. If the patient has lymphoma in the thyroid, the prediction result is 1, otherwise it is 0.

[0055] Compared with the prior art, a method for predicting thyroid lymphoma based on machine learning provided by this embodiment extracts various tumor imaging features from the labeled thyroid ultrasound segmentation images, uses zero importance analysis and correlation screening to remove redundancy to obtain the final screening features, trains a thyroid lymphoma prediction model based on machine learning, and uses the prediction model for the thyroid ultrasound segmentation map to be measured, and finally realizes the accurate prediction of thyroid lymphoma. On the one hand, extracting various types of features including first-order, shape, and high-level texture features can comprehensively reflect the hierarchical features of thyroid lymphoma, provide rich input information for the prediction model, and improve the accuracy of the model after redundancy removal processing; on the other hand, the features based on dimensionality reduction and redundancy removal are used for prediction by a random model, which provides model interpretability while reducing the computational cost and improving the model efficiency.

[0056] Embodiment 2 Another specific embodiment of the present invention discloses a device for predicting thyroid lymphoma based on machine learning, including: An image acquisition module for collecting a plurality of ultrasound images including the thyroid region; A segmentation and annotation module for segmenting the lesion area on the thyroid ultrasound image to generate a thyroid ultrasound segmentation image; and also for annotating the lesion type label including lymphoma on the thyroid ultrasound segmentation image; A feature extraction module for extracting the tumor imaging features of the thyroid ultrasound segmentation image to obtain initial features, and sequentially performing zero importance analysis and correlation screening on the initial features to obtain final screening features with high importance and low correlation; and also for extracting features from the thyroid ultrasound segmentation image to be measured with the suspected lesion area segmented to obtain the corresponding final screening features; A prediction module for training a thyroid lymphoma prediction model based on machine learning based on the final screening features and the annotated lesion type labels; and also for predicting thyroid lymphoma using the thyroid lymphoma prediction model based on the corresponding final screening features; A human-computer interaction module for outputting the prediction result using visualization software and interface for the loaded image and feature data.

[0057] Further, performing zero-importance analysis on the initial features by using the feature extraction module, including: Obtaining an original data set based on the thyroid ultrasound segmentation map labeled with lesion type labels, training a random forest model based on the original data set, and calculating the importance score of each initial feature under the condition of the true label; wherein, the true label is the lesion type label corresponding to the lesion area in the thyroid ultrasound segmentation image; Randomly scrambling the lesion type labels to obtain a pseudo-label data set, retraining the random forest model based on the pseudo-label data set, and calculating the importance score of each initial feature under the condition of the pseudo label; Comparing the distribution of the importance scores of each initial feature under the conditions of the true label and the pseudo label to obtain the zero-importance score of the initial feature; Screening out the features with zero-importance scores greater than a preset zeroing threshold among all the initial features to obtain high-importance features.

[0058] Wherein, the system can perform the prediction of thyroid lymphoma according to the machine learning-based thyroid lymphoma prediction method described in any one of the solutions in Embodiment 1. The relevant parts are mutually referred to, and the description is not repeated in this embodiment.

[0059] Compared with the prior art, a machine learning-based thyroid lymphoma prediction device provided in this embodiment, based on an image acquisition module, a segmentation annotation module, a feature extraction module, a prediction module and a human-computer interaction module, proposes a machine learning-based thyroid lymphoma prediction device, realizes the prediction of thyroid lymphoma, provides a powerful decision support tool for clinicians, and realizes the early warning and precise prevention and control of thyroid lymphoma.

[0060] Embodiment 3 Another specific embodiment of the present invention discloses a computing device for machine learning-based thyroid lymphoma prediction, including: A memory, configured to store computer-executable instructions for machine learning-based thyroid lymphoma prediction; A processor, configured to execute the computer-executable instructions and process the execution of the thyroid lymphoma prediction method described in any one of the solutions in Embodiment 1.

[0061] Those skilled in the art can understand that all or part of the processes of implementing the method in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disc, a read-only memory or a random access memory, etc.

[0062] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered within the protection scope of the present invention.

Claims

1. A method for predicting thyroid lymphoma based on machine learning, characterized in that, Including the following steps: For the thyroid ultrasound images with the lesion area segmented, label the lesion type tags including lymphoma to obtain the original dataset; Extract the tumor imaging features of the original dataset to obtain the initial features, and perform zero importance analysis and correlation screening in sequence to obtain the final screening features with high importance and low correlation; Based on the final screening features and the labeled lesion type tags, use machine learning to train a thyroid lymphoma prediction model; Extract features from the thyroid ultrasound segmentation image to be measured with the suspected lesion area segmented to obtain the corresponding final screening features; Based on the corresponding final screening features, use the thyroid lymphoma prediction model to predict thyroid lymphoma.

2. The thyroid lymphoma prediction method based on machine learning according to claim 1, characterized in that Perform zero importance analysis on the initial features, including: Train a random forest model based on the original dataset, and calculate the importance score of each initial feature under the condition of the true label; wherein, the true label is the lesion type tag corresponding to the lesion area in the thyroid ultrasound segmentation image; Randomly scramble the lesion type tags to obtain a pseudo-label dataset, and retrain the random forest model based on the pseudo-label dataset, and calculate the importance score of each initial feature under the condition of the pseudo-label; Compare the distribution of the importance scores of each initial feature under the conditions of the true label and the pseudo-label to obtain the zero importance score of this initial feature; Screen out the features with zero importance scores greater than the preset zeroing threshold among all the initial features to obtain the high-importance features.

3. The method for predicting thyroid lymphoma based on machine learning according to claim 2, wherein Perform correlation screening on the high-importance features to obtain the final screening features, including: Calculate the correlation between each high-importance feature to obtain a correlation matrix; Sort the high-importance features based on the zero importance score; Traverse each high-importance feature in the sorted order and combine the correlation matrix to perform pairwise comparison, and screen out the features with a correlation lower than the preset correlation threshold to obtain the final screening features.

4. The thyroid lymphoma prediction method based on machine learning according to claim 2, wherein, Calculate the importance score of each initial feature using the Gini index, and obtain the zero importance score of the initial feature based on the following formula: , Among them, is the zero importance score of the initial feature i; is the importance score of the initial feature under the condition of the true label; is the importance score of the initial feature under the condition of the pseudo label; P() represents the probability operation.

5. A method for predicting thyroid lymphoma based on machine learning according to claim 1, characterized in that, The main categories of the initial features include first-order features, shape features, and high-level texture features. Based on the final screening features and the labeled lesion type tags, use random forest to train a thyroid lymphoma prediction model.

6. The method for predicting thyroid lymphoma based on machine learning according to Claim 5, wherein, Adopt an automatic sampling technique and use random forest to train the thyroid lymphoma prediction model, including: Based on the final screening features and the labeled lesion type tags, obtain a training dataset, and generate multiple sub-datasets by sampling with replacement from the training dataset; Based on each sub-dataset, independently train the corresponding decision tree in the thyroid lymphoma prediction model based on random forest; Integrate the prediction results of each decision tree based on the following formula to obtain the final prediction result of thyroid lymphoma: , Among them, is the prediction result of the final thyroid lymphoma; is the prediction result of the i-th decision tree; is the total number of decision trees.

7. A method for predicting thyroid lymphoma based on machine learning according to claim 6, wherein When independently training the corresponding decision tree, sample the sample features and labels in the sub-dataset with replacement, and generate the corresponding training subset based on the following formula: , Among them, is the training subset corresponding to the i-th decision tree; respectively represent the training samples independently sampled from the sample feature distribution and the label distribution obtained from the sample feature distribution.

8. A method for predicting thyroid lymphoma based on machine learning according to any one of claims 1-7, characterized in that, Extract the tumor imaging features of the original dataset, including: Use the grayscale histogram to extract the global intensity features of the thyroid ultrasound segmentation images in the original dataset to obtain first-order features; The geometric shape features are extracted using the area and perimeter of the segmented region to obtain the shape features; The gray-level co-occurrence matrix, gray-level dependency matrix and gray-level run-length matrix are used to extract local structural information and obtain high-level texture features.

9. A thyroid lymphoma prediction device based on machine learning, characterized in that, include: An image acquisition module, used for acquiring an ultrasound image including a thyroid region; A segmentation and annotation module is used to segment the lesion area on the thyroid ultrasound image to generate a thyroid ultrasound segmented image; and is also used to annotate the lesion type label including lymphoma on the thyroid ultrasound segmented image; A feature extraction module is used to extract tumor imaging features of the thyroid ultrasound segmentation image to obtain initial features, and to perform zero importance analysis and correlation screening on the initial features in sequence to obtain final screening features with high importance and low correlation; It is also used to extract features from the thyroid ultrasound segmentation image to be tested that segments the suspected lesion area to obtain the corresponding final screening features; A prediction module, used to obtain a thyroid lymphoma prediction model by using machine learning training based on the final screening features and the annotated lesion type labels; and also used to predict thyroid lymphoma by using the thyroid lymphoma prediction model based on the corresponding final screening features; The human-computer interaction module is used to output prediction results for the loaded images and feature data using visualization software and interface.

10. The thyroid lymphoma prediction device based on machine learning according to claim 9, characterized in that, Using the feature extraction module to perform zero importance analysis on the initial features includes: An original data set is obtained based on the thyroid ultrasound segmentation image marked with lesion type labels, a random forest model is trained based on the original data set, and the importance score of each initial feature under the true label condition is calculated; wherein the true label is the lesion type label corresponding to the lesion area in the thyroid ultrasound segmentation image; Randomly scramble the lesion type labels to obtain a pseudo-label data set, and retrain the random forest model based on the pseudo-label data set to calculate the importance score of each initial feature under the pseudo-label condition; Compare the distribution of the importance scores of each initial feature under the real label and pseudo label conditions to obtain the zero importance score of the initial feature; Among all the initial features, the features whose zero importance scores are greater than the preset zeroing threshold are screened out to obtain high-importance features.

Citation Information

Patent Citations

  • Intracranial primary malignant tumor identification method based on machine learning

    CN116152170A

  • Laboratory mouse-oriented intracranial high-risk aneurysm rupture risk analysis method and system

    CN118864382A

  • Drug interaction prediction method based on multi-dimensional features

    CN119230128A

  • Method And Apparatus Of Establishing Image Search Relevance Prediction Model, And Image Search Method And Apparatus

    US20170330054A1

  • Providing QA training data and training a QA model based on implicit relevance feedbacks

    WO2021146003A1