Method and device for filling in clinical features of thyroid nodules and optimizing prediction of benignity and malignancy
By optimizing the binning of thyroid nodule features and the feature importance parameters, and constructing a training sample set, the problem of high missing data rate of thyroid nodule was solved, the accuracy and adaptability of the prediction model were improved, and the credibility of AI-assisted diagnosis was enhanced.
Patent Information
- Application Number
- CN202511135700.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Existing technologies suffer from high feature loss rates when processing thyroid nodule data, leading to decreased model prediction performance. Traditional methods struggle to adapt to multimodal distribution characteristics and imbalanced cross-center samples, affecting the model's generalization ability and accuracy.
The initial binning boundaries are obtained through decision trees, feature importance parameters are calculated, binning boundaries are adjusted, malignancy prevalence within each bin is used as a feature, a training sample set is constructed, the model is optimized to address the missing value problem, and robustness is improved by combining a multi-model evaluation framework.
It improved the accuracy and generalization ability of predicting benign and malignant thyroid nodules, enhanced the model's cross-center adaptability and prediction accuracy in small hospitals, and improved the credibility of AI-assisted diagnosis by combining it with doctors' decision-making.
Smart Images

Figure CN120727291B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of machine learning, and in particular to a thyroid nodule clinical feature filling and benign and malignant prediction optimization method and device. BACKGROUND
[0002] As a common clinical condition, the benign and malignant differentiation of thyroid nodules is crucial for diagnosis and treatment decisions. With the development of medical informatization, a large amount of multi-dimensional diagnosis and treatment data containing ultrasound images, blood biochemical indicators and patient demographic information has been accumulated in clinical practice. However, such data generally have the problem of feature missing, with a missing rate of up to 15%-50%, mainly due to selective implementation of examination items, equipment differences and omission of electronic medical record entry, etc. Systematic missing of key prediction features such as ultrasound calcification features and elastography scores seriously hinders the development of machine learning-based medical auxiliary diagnosis models. Traditional processing methods are prone to significantly degrade the prediction performance of the model.
[0003] The current mainstream method for feature missing in medical clinical data processing has obvious limitations. Traditional statistical filling methods such as mean / median filling ignore the nonlinear correlation between features, for example, the threshold effect of thyroid globulin level on the probability of nodule malignancy, and simple mean filling will distort the feature distribution; multiple imputation methods improve robustness by constructing multiple filling models, but the Markov chain Monte Carlo method based on them has an exponential growth in computational complexity for high-dimensional features, and there is a significant performance bottleneck on thyroid data sets containing 20 or more high-dimensional features. Directly deleting samples containing missing values in the positive sample-scarce thyroid malignant nodule detection scenario easily leads to the loss of key case data. Experiments show that when the feature missing rate is >20%, this method will reduce the training set size by more than 40%, significantly reducing the model sensitivity. Fixed threshold division of equal-width / equal-frequency binning is difficult to adapt to the multimodal distribution characteristics of thyroid features, such as the significant regional specificity of the interaction effect of ultrasound TI-RADS classification and nodule size, and fixed binning will destroy the collaborative prediction signal between features. In addition, existing researches focus on the filling effect verification of a single prediction model, ignoring the differences in feature sensitivity of different models, for example, the tolerance of neural networks to continuous feature filling errors is significantly lower than that of tree models, and existing methods lack a systematic evaluation framework for multi-model robustness, making it difficult to adapt to clinical multi-model application scenarios. SUMMARY
[0004] The embodiment of the application provides a thyroid nodule clinical feature filling and benign and malignant prediction optimization method and device, uses importance parameters of different thyroid nodule clinical features to perform optimal binning, and uses the prevalence of malignant thyroid nodules in each bin as a bin element feature to train a preset nodule prediction architecture, so that the data complexity is reduced, the key information is highlighted, the missing value problem is solved, and the model generalization capability is enhanced, and the malignant thyroid nodules can be more accurately predicted.
[0005] In a first aspect, the embodiment of the application provides a thyroid nodule clinical feature filling and benign and malignant prediction optimization method, which comprises the following steps:
[0006] An initial binning boundary of each kind of thyroid nodule clinical feature is obtained based on a decision tree, and each kind of thyroid clinical feature is binned using the corresponding initial binning boundary to obtain an initial binning result;
[0007] An importance parameter of each kind of thyroid nodule clinical feature is calculated based on the corresponding initial binning result, an optimal binning boundary of each kind of thyroid nodule clinical feature is adjusted based on the corresponding importance parameter and the initial binning result, and each kind of thyroid nodule clinical feature is re-binned using the corresponding optimal binning boundary to obtain an optimal binning result, wherein the importance parameter is positively correlated with the binning granularity;
[0008] The prevalence of malignant thyroid nodules in each bin is used as a bin element feature, bin element features of all kinds of thyroid nodule clinical features of multiple patients are obtained, and the missing bin element features of the patients are supplemented to serve as a training sample set, a preset nodule prediction architecture is trained using the training sample set to obtain a nodule prediction model;
[0009] Bin element features of multiple thyroid nodule clinical features of a to-be-tested patient are obtained as to-be-tested input features, the to-be-tested input features are input into the nodule prediction model to obtain the prevalence probability of malignant thyroid nodules of the to-be-tested patient.
[0010] In a second aspect, the embodiment of the application provides a thyroid nodule clinical feature filling and benign and malignant prediction optimization device, which comprises the following steps:
[0011] An initial binning module is configured to obtain an initial binning boundary of each kind of thyroid nodule clinical feature based on a decision tree, and bin each kind of thyroid clinical feature using the corresponding initial binning boundary to obtain an initial binning result;
[0012] The binning optimization module calculates an importance parameter of each kind of thyroid nodule clinical feature based on the corresponding initial binning result, adjusts an optimal binning boundary of each kind of thyroid nodule clinical feature based on the corresponding importance parameter and the initial binning result, and re-bins each kind of thyroid nodule clinical feature using the corresponding optimal binning boundary to obtain an optimal binning result, wherein the importance parameter is positively correlated with binning granularity.
[0013] The training module is configured to obtain binning element features of all kinds of thyroid nodule clinical features of a plurality of patients by taking a prevalence rate of malignant thyroid nodules in each bin as a binning element feature, and supplement the binning element features missing in the patients to serve as a training sample set, and train a preset nodule prediction architecture using the training sample set to obtain a nodule prediction model.
[0014] The prediction module is configured to obtain binning element features of a plurality of thyroid nodule clinical features of a to-be-tested patient as to-be-tested input features, and input the to-be-tested input features into the nodule prediction model to obtain a prevalence probability of malignant thyroid nodules of the to-be-tested patient.
[0015] In a third aspect, an electronic device is provided, including a memory and a processor, the memory storing a computer program, and the processor is configured to run the computer program to perform the method for filling in thyroid nodule clinical features and optimizing benign and malignant prediction.
[0016] In a fourth aspect, a readable storage medium is provided, the readable storage medium storing a computer program, the computer program including program codes for controlling a process to perform the method for filling in thyroid nodule clinical features and optimizing benign and malignant prediction.
[0017] The main contributions and innovations of the present application are as follows:
[0018] The embodiments of the present application automatically adjust the binning granularity according to the feature importance, simultaneously complete feature discretization and standardization, retain the diagnostic value of nonlinear features, improve the accuracy and generalization of thyroid nodule benign and malignant prediction without increasing the computational complexity; the embodiments of the present application break through the traditional single-scene single-evaluation quantization path, promote cross-center feature standardization, dynamically select the standardization technology according to the center sample size and feature missing distribution, construct a model sensitivity matrix to quantify the robustness, combine related means to improve the prediction performance of low-quality center data, and guarantee the prediction accuracy of small hospitals; the embodiments of the present application integrate the model prediction result with the doctor's clinical experience decision, integrate the hierarchical interpretable module and clinical rules into medical clinical modeling, and improve the doctor's trust and AI assisted decision-making efficiency.
[0019] The details of one or more embodiments of the application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the application will be apparent from the description of illustrative embodiments of the application and from the claims. BRIEF DESCRIPTION OF DRAWINGS
[0020] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this application, illustrate embodiments of the application and serve to explain the principles of the application. In the drawings:
[0021] Figure 1 is a flow chart of a thyroid nodule clinical feature filling and benign and malignant prediction optimization method according to an embodiment of the application;
[0022] Figure 2 is a structural block diagram of a thyroid nodule clinical feature filling and benign and malignant prediction optimization device according to an embodiment of the application;
[0023] Figure 3 is a hardware structure schematic diagram of an electronic device according to an embodiment of the application. DETAILED DESCRIPTION
[0024] The illustrative examples set forth in the following description of embodiments of the application are provided to illustrate the application and to provide a comprehensive description of various features and embodiments of the application. The examples do not represent an exhaustive list of all possible embodiments of the application. Rather, the examples are provided to illustrate the application and to provide a comprehensive description of various features and embodiments of the application.
[0025] It should be noted that the steps of the methods in other embodiments are not necessarily performed in the order shown and described in this specification. In some other embodiments, the steps of the methods can include more or fewer steps than those described in this specification. Furthermore, individual steps described in this specification can be split into multiple steps to be described in other embodiments; and multiple steps described in this specification can be combined into a single step to be described in other embodiments.
[0026] Embodiment one
[0027] The embodiment of the application provides a thyroid nodule clinical feature filling and benign and malignant prediction optimization method, which uses the importance parameters of different thyroid nodule clinical features to perform optimal binning, and uses the malignant thyroid nodule prevalence rate in each bin as a bin element feature to train a preset nodule prediction architecture, so as to reduce data complexity, highlight key information, cope with missing value problems and enhance model generalization ability, and more accurately predict malignant thyroid nodules. Specifically, referenceFigure 1 The method comprises:
[0028] An initial binning boundary of each thyroid nodule clinical feature is obtained based on a decision tree, and each thyroid clinical feature is binned using the corresponding initial binning boundary to obtain an initial binning result.
[0029] An importance parameter of each thyroid nodule clinical feature is calculated based on the corresponding initial binning result, the optimal binning boundary of each thyroid nodule clinical feature is adjusted based on the corresponding importance parameter and the initial binning result, and each thyroid nodule clinical feature is re-binned using the corresponding optimal binning boundary to obtain an optimal binning result, wherein the importance parameter is positively correlated with the binning granularity.
[0030] The binning element features of all kinds of thyroid nodule clinical features of a plurality of patients are obtained by taking the prevalence of malignant thyroid nodules in each bin as a binning element feature, and the missing binning element features of the patients are supplemented as a training sample set, and a preset nodule prediction architecture is trained using the training sample set to obtain a nodule prediction model.
[0031] The binning element features of a plurality of thyroid nodule clinical features of a patient to be tested are obtained as input features to be tested, and the input features to be tested are input into the nodule prediction model to obtain the prevalence probability of malignant thyroid nodules of the patient to be tested.
[0032] In some embodiments, the initial binning boundary of each thyroid nodule clinical feature is obtained by setting the maximum depth of the decision tree.
[0033] Specifically, before binning, each thyroid nodule clinical feature is arranged in order of size, and each thyroid nodule clinical feature is binned based on the corresponding initial binning boundary to obtain an initial binning result.
[0034] In some specific embodiments, the thyroid nodule clinical features include calcification features, aspect ratio features, echo features, boundary features, and blood flow features.
[0035] Specifically, the calcification features refer to the micro-calcification foci that may be produced by cancer cells in the metabolic process, so the appearance of micro-calcification under ultrasonic imaging often indicates a higher risk of malignant thyroid nodules. In contrast, coarse calcification has larger calcification foci and relatively irregular shape, and coarse calcification may also appear in benign nodules.
[0036] Specifically, the aspect ratio refers to the ratio of the anteroposterior diameter to the transverse diameter of the nodule in the ultrasound image. When the aspect ratio is greater than 1, the nodule appears to be vertically long, which often indicates an increased likelihood of malignancy. Benign nodules tend to be horizontally long, with an aspect ratio less than 1, and have a relatively flat and round overall shape.
[0037] Specifically, the echo feature is one of the key observation points in thyroid nodule ultrasound examination. Low echo nodules have a relatively high risk of malignancy, while high echo nodules are more commonly seen in benign lesions.
[0038] Specifically, the degree of boundary clarity of the nodule is crucial in determining its nature. Most benign nodules have clear boundaries and can be clearly distinguished from the surrounding thyroid tissue. The edge is smooth, indicating that the nodule grows relatively regularly without irregular infiltration of the surrounding tissue. The boundary of malignant nodules is often unclear, often appearing blurred, and the boundary with the surrounding tissue is not clear.
[0039] Specifically, observing the blood flow signals within and around the nodule through ultrasound is also helpful in determining the nature of the nodule. The blood flow signals of benign nodules are usually relatively regular and uniform, often appearing as a peripheral ring. Malignant nodules, due to their active cell proliferation, require more nutrients, often have abundant and chaotic blood flow signals, not only an increase in peripheral blood flow, but also more branched and penetrating blood flow distribution inside.
[0040] In some specific embodiments, the importance score of each bin in the initial binning result is calculated respectively, and the importance scores of each bin under the same thyroid nodule clinical feature are integrated to obtain the importance parameter of each thyroid nodule clinical feature.
[0041] Specifically, the IV (information value) is used to calculate the importance score of each bin, and the formula of the importance parameter is as follows:
[0042]
[0043] wherein, is the importance parameter of the thyroid nodule clinical feature , the importance score of bin i, and k is the number of bins of the thyroid nodule clinical feature .
[0044] Further, the information gain of the thyroid nodule clinical feature under the current bin boundary is calculated, the information gain and the importance parameter are weighted and summed to obtain the information gain feature of the corresponding thyroid nodule clinical feature, and the current bin boundary when the information gain feature is maximum is taken as the optimal bin boundary of the corresponding thyroid nodule clinical feature.
[0045] It is worth mentioning that the maximum value of the information gain feature is found by traversing and iterating. In the first round of traversal iteration, the current bin boundary is the initial bin boundary. In subsequent iterations, the current bin boundary is continuously updated until the optimal bin boundary is obtained.
[0046] Specifically, the formula for obtaining the optimal bin boundaries is illustrated as follows:
[0047]
[0048] in, and This is a balancing coefficient used to balance the influence of information gain and importance parameters. Clinical features of thyroid nodules Information gain function, Clinical features of thyroid nodules Importance parameter, B is the current bin boundary, B={ }
[0049] Specifically, this scheme comprehensively considers information gain and importance parameters, and uses a balance coefficient to balance their influence. This avoids the one-sidedness of relying solely on information gain or feature importance as a single factor to select binning boundaries. It can more scientifically and comprehensively measure the value of clinical features of thyroid nodules under different binning conditions, and make the selected binning boundaries more in line with actual application needs.
[0050] Furthermore, in this scheme, the importance parameter of the clinical features of thyroid nodules is positively correlated with the binning granularity. That is, the more important the clinical features of thyroid nodules are for the determination of benign or malignant nature, the finer the binning granularity will be.
[0051] In some specific embodiments, the number of patients with malignant thyroid nodules in each compartment is obtained, and the ratio of the number of patients with malignant thyroid nodules in the compartment to the total number of patients in the compartment is used as the compartment element feature.
[0052] Specifically, the formula for binning element features is expressed as follows:
[0053]
[0054] in, , where m is the number of patients with malignant thyroid nodules in the bin, and N is the total number of patients in the bin.
[0055] Specifically, the data used in this program during training was retrieved from various hospital centers, so the benign or malignant nature of each patient's thyroid nodules was diagnosed and marked by the doctor.
[0056] In some embodiments, due to the variety of thyroid clinical features and the complexity of the data, directly using the original data as the input of the model will cause a large deviation in the prediction result, and overfitting is prone to occur. The present scheme uses the prevalence of malignant thyroid nodules in each bin as the bin element feature, and performs an induction and simplification process on the original complex clinical feature data. Integrating the originally scattered and complex data into the relatively unified index of prevalence corresponding to different bins makes the input data form more regular and simple, which helps to reduce the complexity of the data and enables the subsequent model training process to focus more efficiently on key information, avoiding overfitting due to overly complex data.
[0057] In addition, the present scheme focuses on the prevalence of malignant thyroid nodules in each bin, which is directly related to the key goal of distinguishing between benign and malignant thyroid nodules. In this way, the relationship between different clinical features and the malignancy of nodules in different bins can be highlighted, making it easier for the model to capture feature patterns that are important for distinguishing between benign and malignant nodules during training, thereby strengthening the model's ability to distinguish between benign and malignant nodules and more accurately extracting effective prediction information hidden in numerous clinical features.
[0058] In some embodiments, missing patients with the same type of bin element feature are placed in a set as a missing patient set, and the ratio of the number of malignant thyroid nodule patients in the missing patient set to the number of all missing patients in the bin is used as a supplementary bin element feature. The corresponding type of bin element feature for each missing patient in the corresponding missing patient set is completed using the supplementary bin element feature.
[0059] For example, for the age feature, all patients with missing age are placed in a set as an age missing patient set, and then the prevalence of the number of malignant thyroid nodule patients in the age missing patient set is used as a supplementary bin element feature. Assuming it is 0.6, then the age bin element feature in the age missing patient set is supplemented to 0.6.
[0060] In some embodiments, during the training of the nodule prediction architecture, the weight of the misjudgment sample is dynamically adjusted, and the weight of the misjudgment sample is introduced into the prediction loss function of the nodule prediction architecture. The formula for adjusting the weight of the misjudgment sample is as follows:
[0061]
[0062] wherein, is the weight of the i-th sample in the training sample set in the t+1 iteration, is the weight of the i-th sample in the training sample set in the t iteration, is a prediction loss function, is a weight of the i-th sample in the t-th iteration of the training sample set, is a prediction result of the i-th sample. is a weight adjustment coefficient, , is a weighted error rate of the t-th round model, is a compensation hyperparameter.
[0063] The formula for introducing the weight of the misjudgment sample into the prediction loss function of the nodule prediction architecture is as follows:
[0064]
[0065] wherein, is a prediction loss function, is a weight of the i-th sample in the t-th iteration of the training sample set, is a prediction result of the i-th sample. is a prediction result of the i-th sample.
[0066] Specifically, by dynamically adjusting the weight of the misjudgment sample and introducing the weight of the misjudgment sample into the prediction loss function of the nodule prediction architecture, the model can be guided to focus on the learning of the misjudgment sample, prompting the model to adjust its parameters in the direction of more accurately processing such easily misjudged cases, and ultimately achieving the effect of improving the accuracy and generalization ability of the model.
[0067] In some specific embodiments, multiple types of nodule prediction architectures are constructed, and the training data set is used to train the multiple types of nodule prediction architectures to obtain multiple types of nodule prediction models. A prediction sensitivity matrix is constructed for each type of nodule prediction model, which is used to evaluate the prediction accuracy of the corresponding nodule prediction model under different feature missing rates. The feature missing rate represents the missing situation of the clinical features of thyroid nodules. According to the feature missing rate in the database of different hospital centers, the nodule prediction model with the highest prediction accuracy is selected to predict the probability of malignant thyroid nodule disease of the patient to be tested.
[0068] Specifically, due to different medical habits of different hospitals, the project inspection of thyroid nodules will be different. Some hospitals will check all the clinical features of thyroid nodules including calcification features, aspect ratio features, echo features, boundary features, and blood flow features, while some hospitals may only check calcification features, aspect ratio features, echo features, and boundary features. For nodule prediction models, when there is an uncertain feature missing, different types of models may be affected differently. For example, the XGboost model has a prediction accuracy of 0.95 for a feature with a missing rate of 0, but a prediction accuracy of 0.78 for a feature with a missing rate of 30%. In the logistic regression model, the prediction accuracy for a feature with a missing rate of 0 is 0.91, and the prediction accuracy for a feature with a missing rate of 30% is 0.85.
[0069] Further, the sensitivity matrix is denoted as , and the element in the sensitivity matrix is , denotes the relative performance value of the index vector q of the nodule prediction model m at the missing rate k.
[0070] Specifically, denotes the index vector of the nodule prediction model m on the dataset with the missing rate of 0, denotes the index vector of the nodule prediction model m on the dataset with the missing rate of k.
[0071] Specifically, The calculation formula of is:
[0072]
[0073] wherein, is the index of the nodule prediction model m on the dataset with the missing rate of k for prediction, wherein:
[0074] is the sensitivity of the nodule prediction model m on the dataset with the missing rate of k for prediction, and the evaluation is the proportion of the true positive (malignant nodule) samples in the dataset that are predicted as positive by the model, and the calculation formula is:
[0075]
[0076] wherein, represents the number of samples with the true label as positive and the prediction as positive, and represents the number of samples with the true label as positive but the prediction as negative. The sensitivity represents the recall rate of the positive samples, and if the positive is not correctly predicted, it is the “missed diagnosis” case.
[0077] is the specificity of the nodule prediction model m on the dataset with the missing rate of k for prediction, and the evaluation is the proportion of the true negative (benign nodule) samples that are predicted as negative by the model, and the calculation formula is:
[0078]
[0079] wherein, represents the number of samples with the true label as negative and the prediction as negative, and represents the number of samples with the true label as negative but the prediction as positive. The specificity represents the recall rate of the negative samples, and if the negative is not correctly predicted, it is the “misdiagnosis” case.
[0080] is the specificity of the nodule prediction model m on the dataset with the missing rate of k The area under the ROC curve and the false positive rate (AUC) is used for prediction. It assesses the model's ability to distinguish between positive (positive / malignant) and negative (negative / benign) classes. A larger AUC indicates a stronger discriminatory ability. The formula for calculating AUC is:
[0081]
[0082]
[0083]
[0084]
[0085] Where, represents the number of patient samples with positive thyroid nodules in the dataset, represents the number of patient samples with negative thyroid nodules in the dataset, represents the output score of the prediction model, and represents the indicator function that consists of a pair of negative and positive samples.
[0086] For nodule prediction model m on a dataset with a missing rate of k The accuracy of the predictions is used to evaluate the proportion of all correctly predicted samples out of the total sample size. The calculation formula is:
[0087]
[0088] For nodule prediction model m on a dataset with a missing rate of k Prediction accuracy when performing predictions ( The harmonic mean of positive samples and recall is a comprehensive evaluation metric, and its calculation formula is as follows:
[0089]
[0090]
[0091]
[0092] For nodule prediction model m on a dataset with a missing rate of k The positive predictive value used in the prediction is equivalent to precision, which assesses the proportion of samples predicted to be positive that are actually positive. Its calculation formula is as follows:
[0093]
[0094] For nodule prediction model m on a dataset with a missing rate of k The negative predictive value of the prediction on the upper part is evaluated as the proportion of the samples with true label negative in the samples predicted as negative, and the calculation formula is:
[0095]
[0096] In some specific embodiments, the influence of each bin element feature of the patient to be tested on the predicted probability of malignant thyroid nodule is quantified by using an explainable method, a waterfall chart is generated to intuitively show how each feature affects the probability of malignancy (such as microcalcification contributes +23%), and a decision tree is used to visualize the key judgment path (such as “aspect ratio > 1 → malignant probability increases by 18%”).
[0097] Further, the generated waterfall chart is combined with the thyroid clinical guidelines to generate a structured report template, and the clinician modifies or confirms the probability of malignant thyroid nodule based on the structured report template.
[0098] Embodiment two
[0099] Based on the same concept, referring to Figure 2 The application also proposes a thyroid nodule clinical feature filling and malignant and benign prediction optimization device, which comprises:
[0100] An initial binning module obtains initial binning boundaries of each kind of thyroid nodule clinical feature based on a decision tree, and bins each kind of thyroid clinical feature using the corresponding initial binning boundaries to obtain an initial binning result;
[0101] A binning optimization module calculates an importance parameter of each kind of thyroid nodule clinical feature based on the corresponding initial binning result, adjusts the optimal binning boundary of each kind of thyroid nodule clinical feature based on the corresponding importance parameter and the initial binning result, and re-bins each kind of thyroid nodule clinical feature using the corresponding optimal binning boundary to obtain an optimal binning result, wherein the importance parameter is positively correlated with the binning granularity;
[0102] A training module is configured to use the prevalence of malignant thyroid nodule in each bin as a bin element feature, obtain bin element features of all kinds of thyroid nodule clinical features of multiple patients, and supplement the missing bin element features of the patients to serve as a training sample set, and train a predetermined nodule prediction architecture using the training sample set to obtain a nodule prediction model;
[0103] A prediction module is configured to obtain bin element features of multiple thyroid nodule clinical features of a patient to be tested as a to-be-tested input feature, input the to-be-tested input feature into the nodule prediction model to obtain the prevalence of malignant thyroid nodule of the patient to be tested.
[0104] Embodiment three
[0105] The embodiment also provides an electronic device, referring to Figure 3 comprising a memory 404 and a processor 402, the memory 404 storing a computer program, and the processor 402 being configured to execute the computer program to perform the steps in any of the above method embodiments.
[0106] Specifically, the processor 402 described above can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the application.
[0107] The memory 404 can include a mass storage that stores data or instructions. For example, and without limitation, the memory 404 can include a Hard Disk Drive (HDD), a floppy disk drive, a Solid State Drive (SSD), a flash drive, a Compact Disc Read Only Memory (CD-ROM), a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. The memory 404 can be removable and / or non-removable (or fixed) as appropriate. The memory 404 can be internal or external as appropriate. In particular embodiments, the memory 404 is a Non-Volatile memory. In particular embodiments, the memory 404 includes a Read-Only Memory (ROM) and a Random Access Memory (RAM). The ROM can be a mask-programmed ROM, a Programmable ROM (PROM), an Erasable PROM (EPROM), an Electrically EPROM (EEPROM), an Electrically Alterable ROM (EAROM), or a FLASH memory, or a combination of two or more of these, as appropriate. The RAM can be a Static Random-Access Memory (SRAM) or a Dynamic Random Access Memory (DRAM), which can be a Fast Page Mode Dynamic Random Access Memory (FPMDRAM), an Extended Data Output Dynamic Random Access Memory (EDODRAM), a Synchronous Dynamic Random-Access Memory (SDRAM), or the like, as appropriate.
[0108] The memory 404 can be used to store or cache various data files needed for processing and / or communication, and possible computer program instructions executed by the processor 402.
[0109] The processor 402 implements the above-mentioned thyroid nodule clinical feature imputation and benign / malignant prediction optimization method of any one of the embodiments by reading and executing the computer program instructions stored in the memory 404.
[0110] Optionally, the above-mentioned electronic device can further include a transmission device 406 connected to the processor 402 and an input / output device 408 connected to the processor 402.
[0111] The transmission device 406 can be used to receive or send data via a network. Specific examples of the network can include a wired or wireless network provided by a communication provider of the electronic device. In one example, the transmission device includes a network adapter (NIC) that can be connected to other network devices through a base station to communicate with the Internet. In one example, the transmission device 406 can be a radio frequency (RF) module for communicating with the Internet in a wireless manner.
[0112] The input / output device 408 is used for inputting or outputting information. In the present embodiment, the input information can be various thyroid nodule clinical features, and the output information can be the probability of suffering from a malignant thyroid nodule.
[0113] Optionally, in the present embodiment, the processor 402 can be configured to perform the following steps by computer programs:
[0114] obtain initial binning boundaries of each thyroid nodule clinical feature based on the decision tree, and bin each thyroid clinical feature using the corresponding initial binning boundaries to obtain an initial binning result;
[0115] calculate an importance parameter of each thyroid nodule clinical feature based on the corresponding initial binning result, adjust the optimal binning boundaries of each thyroid nodule clinical feature based on the corresponding importance parameter and the initial binning result, and re-bin each thyroid nodule clinical feature using the corresponding optimal binning boundaries to obtain an optimal binning result, wherein the importance parameter is positively correlated with the binning granularity;
[0116] The binning element features of all kinds of thyroid nodules clinical features of multiple patients are obtained by taking the prevalence of malignant thyroid nodules in each bin as the binning element feature, and the missing binning element features of the patients are supplemented as a training sample set, and a preset nodules prediction architecture is trained by using the training sample set to obtain a nodules prediction model.
[0117] The binning element features of multiple thyroid nodules clinical features of the to-be-tested patient are obtained as to-be-tested input features, and the to-be-tested input features are input into the nodules prediction model to obtain the prevalence probability of malignant thyroid nodules of the to-be-tested patient.
[0118] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and this embodiment will not be repeated here.
[0119] Generally, various embodiments can be implemented in hardware or special-purpose circuitry, software, logic or any combination thereof. Some aspects of the application can be implemented in hardware, while other aspects can be implemented by firmware or software executed by a controller, microprocessor or other computing device, but the application is not limited thereto. Although various aspects of the application can be illustrated and described as a block diagram, flow chart, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques or methods described herein can be implemented in hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0120] Embodiments of the application can be implemented by computer software executable by a data processor of the mobile device such as in the processor entity, or by hardware, or by a combination of software and hardware. Computer software or programs, also called program products when executed, including software routines, applets, and / or macros, can be stored in any apparatus-readable data storage medium and they include program instructions to implement specific tasks. The program (which can be a component of a program product) can include one or more computer-executable components such as the components shown in the flow charts of FIGS. 1-4, and / or program instructions that, when executed by one or more processors, perform the steps described herein. The one or more computer-executable components can be one or more Figure 3 Any block in the logic flow of the method described herein, such as in FIGS. 1-4, can represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). In some embodiments, the functions noted in the blocks can occur out of the order as shown in the figures. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved.
[0121] Those skilled in the art should understand that each technical feature of the above embodiments can be combined arbitrarily, and for the sake of brevity, each technical feature in the above embodiments is not described in all possible combinations, however, as long as the combination of the technical features does not exist, it should be considered as the scope of the description.
[0122] The above embodiments only express several implementation manners of the present application, the description is more specific and detailed, but it should not be understood as the limitation of the scope of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method for filling and optimizing the prediction of benign and malignant of thyroid nodule clinical features, characterized in that, The method comprises the following steps: An initial binning boundary of each thyroid nodule clinical feature is obtained based on a decision tree, and each thyroid clinical feature is binned using the corresponding initial binning boundary to obtain an initial binning result. An importance parameter of each thyroid nodule clinical feature is calculated based on the corresponding initial binning result, the optimal binning boundary of each thyroid nodule clinical feature is adjusted based on the corresponding importance parameter and the initial binning result, and each thyroid nodule clinical feature is re-binned using the corresponding optimal binning boundary to obtain an optimal binning result, wherein the importance parameter is positively correlated with the binning granularity. The prevalence of malignant thyroid nodules in each bin is taken as a bin element feature, bin element features of all kinds of thyroid nodule clinical features of multiple patients are obtained, and missing bin element features of the patients are supplemented to serve as a training sample set, a preset nodule prediction architecture is trained using the training sample set to obtain a nodule prediction model, wherein a patient with missing bin element features is taken as a missing patient, missing patients with missing bin element features of the same kind are taken into a set as a missing patient set, a ratio of the number of malignant thyroid nodule patients in the missing patient set to the number of all missing patients in the bin is taken as a supplemented bin element feature, and the supplemented bin element feature is used to complete the corresponding kind of bin element feature of each missing patient in the corresponding missing patient set. Bin element features of multiple thyroid nodule clinical features of a to-be-tested patient are obtained as to-be-tested input features, and the to-be-tested input features are input into the nodule prediction model to obtain a prevalence probability of malignant thyroid nodules of the to-be-tested patient.
2. The method of claim 1, wherein the method is characterized by, Importance scores of each bin in the initial binning result are calculated respectively, and the importance scores of each bin under the same kind of thyroid nodule clinical feature are integrated to obtain an importance parameter of each kind of thyroid nodule clinical feature. 3.The thyroid nodule clinical feature filling and benign-malignant prediction optimization method according to claim 1, characterized in that, An information gain of the thyroid nodule clinical feature under the current binning boundary is calculated, the information gain and the importance parameter are weighted and summed to obtain an information gain feature of the corresponding thyroid nodule clinical feature, and the current binning boundary when the information gain feature is maximum is taken as the optimal binning boundary of the corresponding thyroid nodule clinical feature.
4. The method of claim 1, wherein the method is characterized by, The number of malignant thyroid nodule patients in each bin is obtained, and a ratio of the number of malignant thyroid nodule patients in the bin to the number of all patients in the bin is taken as a bin element feature.
5. The method of claim 1, wherein the method is used for filling in the clinical features of thyroid nodules and predicting the benignity or malignancy of the thyroid nodules. During the training of the nodule prediction architecture, the weight of a misjudgment sample is dynamically adjusted, and the weight of the misjudgment sample is introduced into a prediction loss function of the nodule prediction architecture, and the formula for adjusting the weight of the misjudgment sample is as follows: wherein, is the weight of the i-th sample in the training sample set at the t+1-th iteration, is the weight of the i-th sample in the training sample set at the t-th iteration, is an indicator function, which is 1 when a misjudgment occurs, and 0 otherwise, is a weight adjustment coefficient, , is the weighted error rate of the t-th iteration model, is a compensation hyperparameter; The formula for introducing the weight of the misjudgment sample into the prediction loss function of the nodule prediction architecture is as follows: wherein, is a prediction loss function, is a weight of the i-th sample in the t-th iteration of the training sample set, is a prediction result of the i-th sample.
6. The method of claim 1, wherein the method is used for filling and optimizing the prediction of benign and malignant thyroid nodules based on clinical features. The plurality of types of nodule prediction architectures are constructed, and the plurality of types of nodule prediction architectures are trained simultaneously using the training data set to obtain a plurality of types of nodule prediction models. A prediction sensitivity matrix is constructed for each type of nodule prediction model, which is used to evaluate the prediction accuracy of the corresponding nodule prediction model under different feature missing rates. The feature missing rate represents the missing situation of the thyroid nodule clinical features. The nodule prediction model with the highest prediction accuracy is selected according to the feature missing rate in the database of different hospital centers to predict the probability of malignant thyroid nodule of the patient to be tested.
7. A device for filling and optimizing the prediction of benign and malignant of thyroid nodule clinical features, characterized in that, Comprise: An initial binning module, which obtains an initial binning boundary of each type of thyroid nodule clinical feature based on a decision tree, and bins each type of thyroid clinical feature using the corresponding initial binning boundary to obtain an initial binning result; A binning optimization module, which calculates an importance parameter of each type of thyroid nodule clinical feature based on the corresponding initial binning result, adjusts the optimal binning boundary of each type of thyroid nodule clinical feature based on the corresponding importance parameter and the initial binning result, and rebins each type of thyroid nodule clinical feature using the corresponding optimal binning boundary to obtain an optimal binning result. The importance parameter is positively correlated with the binning granularity; A training module, which uses the incidence of malignant thyroid nodule in each bin as a bin element feature, obtains the bin element features of all types of thyroid nodule clinical features of a plurality of patients, and supplements the missing bin element features of the patients to obtain a training sample set. The training sample set is used to train a predetermined nodule prediction architecture to obtain a nodule prediction model. The patients with missing bin element features are regarded as missing patients. The missing patients with missing bin element features of the same type are put into a set as a missing patient set. The ratio of the number of patients with malignant thyroid nodule in the missing patient set to the number of all missing patients in the bin is used as a supplemented bin element feature. The corresponding type of bin element feature of each missing patient in the corresponding missing patient set is supplemented using the supplemented bin element feature. A prediction module, which obtains the bin element features of the plurality of types of thyroid nodule clinical features of the patient to be tested as test input features, and inputs the test input features into the nodule prediction model to obtain the probability of malignant thyroid nodule of the patient to be tested. 8.An electronic device comprising a memory and a processor, the electronic device comprising: The memory stores a computer program, and the processor is configured to run the computer program to execute the thyroid nodule clinical feature filling and benign and malignant prediction optimization method of any one of claims 1-6.
9. A readable storage medium, characterized by, The readable storage medium stores a computer program, and the computer program includes program code for controlling the process to execute the process. The program code is executed by the processor to implement the thyroid nodule clinical feature filling and benign and malignant prediction optimization method of any one of claims 1-6.
Citation Information
Patent Citations
Advanced nasopharynx cancer treatment effect prediction system based on deep learning
CN119132582A
Multi-label medical image classification method
CN120107692A