Prediction method of papillary thyroid carcinoma Ki-67 expression level based on deep learning
Through deep learning-based methods, feature extraction and combination processing of B-mode ultrasound images of patients with papillary thyroid cancer has been solved, and the problem of being unable to quickly and accurately predict Ki-67 expression levels in the prior art is solved, achieving efficient and accurate prediction results.
Patent Information
- Application Number
- CN202510223664.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art cannot quickly and accurately predict the Ki-67 expression level of papillary thyroid carcinoma, making it difficult to accurately screen out cases with possible proliferation progress in clinical practice.
Using a deep learning-based method, image preprocessing of B-mode ultrasound images of patients with papillary thyroid carcinoma is constructed, imaging microscopy features are extracted, and feature combination processing is performed in combination with clinical features. Finally, input the support vector machine classifier for prediction.
The rapid and accurate prediction of Ki-67 expression level of thyroid papillary carcinoma is achieved, and the accuracy and efficiency of clinical diagnosis are improved.
Smart Images

Figure CN120147725A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the technical field of big data processing, and particularly relates to a prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning. Background Art
[0002] Papillary Thyroid Carcinoma (PTC) is the most common subtype of thyroid cancer, and its incidence has increased rapidly in recent years. Although most PTCs exhibit an indolent biological behavior, there is still a part with rapid proliferation and progression. Although the guidelines recommend active surveillance rather than immediate surgery for low-risk PTCs, how to accurately screen out PTCs that may proliferate and progress and give appropriate intervention in a timely manner is a difficult point faced clinically. The Ki-67 index, as an important indicator of tumor cell proliferation activity, is an independent adverse prognostic factor for the proliferation and progression of PTC, and is of great significance for evaluating the prognosis of patients and guiding individualized treatment plans. Traditional Ki-67 detection methods rely on immunohistochemical staining of tissue samples, which have problems such as invasiveness, time-consuming, and high cost.
[0003] Therefore, developing a rapid and accurate Ki-67 expression level prediction method has important clinical value for guiding the precise clinical treatment of PTC. Summary of the Invention
[0004] Based on this, in view of the defect that the prior art cannot rapidly and accurately predict the Ki-67 expression level, it is necessary to provide a prediction method, device, storage medium, and electronic device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning.
[0005] In a first aspect, an embodiment of the present invention provides a prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning, the method comprising:
[0006] Performing image preprocessing on any one of a plurality of B-mode ultrasound images of a target object to obtain a corresponding preprocessed image, where the target object is a patient with papillary thyroid carcinoma;
[0007] Constructing a preset deep learning model;
[0008] Inputting the preprocessed image into the preset deep learning model for feature extraction to obtain target radiomics features, where the target radiomics features are radiomics features that can reflect the Ki-67 expression level of the papillary thyroid carcinoma of the target object;
[0009] Obtaining the processed clinical features of the target object and obtaining the target radiomics features;
[0010] Perform feature combination processing on the processed clinical features and the target radiomics features to obtain corresponding combined features;
[0011] Input the combined features into a classifier of a support vector machine for prediction, and output a prediction result, where the prediction result can predict the Ki-67 expression level of papillary thyroid carcinoma in the target object.
[0012] Optionally, the components of the preset deep learning model include: using the ResNet-50 model as the basic model architecture, integrating multiple target modules, where the multiple target modules include the SE module, the CBAM module, the FPN module, and the clinical attention module. The SE module is used to re-weight the channels of the feature map, the CBAM module is used to enhance the feature extraction performance, the FPN module is used to construct a multi-scale feature pyramid, and the clinical attention module is used to enhance the interaction process between the image features and the clinical features.
[0013] Optionally, constructing the preset deep learning model includes:
[0014] Using the ResNet-50 model as the basic model architecture, constructing a corresponding SE-ResNet50 model through the SE module, and replacing each residual block of the SE-ResNet50 model with a corresponding SE module version;
[0015] On the basis of the SE-ResNet50 model, fuse the CBAM module and the FPN module;
[0016] Define the flow mode of data through the network through forward propagation;
[0017] Enhance the interaction process between the image features and the clinical features through the clinical attention module.
[0018] Optionally, the process of enhancing the interaction process between the image features and the clinical features through the clinical attention module includes:
[0019] Map the image features to the clinical feature space and map the clinical features to the image feature space to implement the processing process of feature space mapping;
[0020] Calculate the attention weights through the ReLU activation function and the Sigmoid activation function;
[0021] Based on the attention weights, perform weighted processing on the image features and the clinical features to enhance the first representation ability of the image features and the second representation ability of the clinical features.
[0022] Optionally, on the basis of the SE-ResNet50 model, fusing the CBAM module and the FPN module includes:
[0023] Obtain the SE-ResNet50 model and remove the original fully connected layer of the SE-ResNet50 model;
[0024] Add an FPN module, which extracts features from multiple different stages of the SE-ResNet50 model and constructs a multi-scale feature pyramid through upsampling and weighted summation;
[0025] For the multiple different stages, add a plurality of CBAM modules that match the number of stages to enhance the channel and spatial attention of the feature map through the CBAM modules, wherein each level of the feature pyramid corresponds to one CBAM module;
[0026] For the feature channels of each level of the feature pyramid, perform global average pooling and flattening, and perform classification processing through a linear classifier.
[0027] Optionally, using the ResNet-50 model as the basic model architecture and constructing the corresponding SE-ResNet50 model through the SE module includes:
[0028] Compress the spatial dimension in a global average pooling manner through the SE module;
[0029] Generate the weight of each channel through the SE module:
[0030] Generate the weight of each channel through a first fully connected layer, a second fully connected layer, and a Sigmoid activation function. The first fully connected layer is used to reduce the number of channels, and the second fully connected layer is used to restore to the original number of channels.
[0031] Optionally, the method further includes:
[0032] Set the hyperparameters of the preset deep learning model;
[0033] The setting of the hyperparameters of the preset deep learning model includes:
[0034] Set the maximum number of iterations to 600;
[0035] Set the learning rate to 2e-5;
[0036] Set the optimizer to the Adam optimizer;
[0037] Set the decay weight to an L2 regularization term of 1e-3 to control the magnitude of the parameters through the decay weight to prevent overfitting;
[0038] Set the updated batch size to 32.
[0039] Optionally, the preset deep learning model adopts a cross-entropy loss function, and the cross-entropy loss function is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label; or,
[0040] The preset deep learning model adopts a learning rate decay strategy; and,
[0041] The preset deep learning model adopts a ReduceLROnPlateau scheduler, so that through the ReduceLROnPlateau scheduler, if the loss does not improve within a preset number of consecutive cycles, the corresponding learning rate will decay to a preset value.
[0042] Optionally, the image preprocessing of any one of the multiple B-mode ultrasound images of the target object includes:
[0043] Randomly select any one of the multiple B-mode ultrasound images of the target object as the current image;
[0044] Perform the following preprocessing operations on the current image:
[0045] Perform preprocessing of image enhancement on the current image using histogram equalization and Canny operator to enhance the texture information and edge information of the corresponding image;
[0046] Perform image data augmentation processing on the current image using at least one of image data augmentation methods including random cropping, random horizontal flipping, and normalization processing to obtain an image after augmentation processing;
[0047] Adopt a channel fusion method to combine the current image and the image after augmentation processing for feature extraction;
[0048] Traverse the remaining images in the multiple B-mode ultrasound images of the target object except the current image, and perform preprocessing operations on any one of the remaining images.
[0049] In a second aspect, an embodiment of the present invention provides a prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning, and the device includes:
[0050] A preprocessing module, configured to perform image preprocessing on any one of the multiple B-mode ultrasound images of a target object to obtain a corresponding preprocessed image, where the target object is a patient with papillary thyroid carcinoma;
[0051] A model construction module for constructing a preset deep learning model;
[0052] A feature extraction module for inputting the preprocessed image into the preset deep learning model for feature extraction to obtain target radiomics features, where the target radiological features are radiomics features that can reflect the Ki-67 expression level of papillary thyroid carcinoma of the target object;
[0053] An acquisition module for acquiring the processed clinical features of the target object and acquiring the target radiomics features;
[0054] A feature combination module for performing feature combination processing on the processed clinical features and the target radiomics features to obtain corresponding combined features;
[0055] A prediction module for inputting the combined features into a classifier of a support vector machine for prediction and outputting a prediction result, where the prediction result can predict the Ki-67 expression level of papillary thyroid carcinoma of the target object.
[0056] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method of the first aspect.
[0057] In a fourth aspect, an electronic device is provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method of the first aspect is implemented.
[0058] In an embodiment of the present invention, any one of a plurality of B-mode ultrasound images of a target object is subjected to image preprocessing to obtain a corresponding preprocessed image, where the target object is a patient with papillary thyroid carcinoma; a preset deep learning model is constructed; the preprocessed image is input into the preset deep learning model for feature extraction to obtain target radiomics features, where the target radiological features are radiomics features that can reflect the Ki-67 expression level of papillary thyroid carcinoma of the target object; the processed clinical features of the target object are acquired, and the target radiomics features are acquired; the processed clinical features and the target radiomics features are subjected to feature combination processing to obtain corresponding combined features; and the combined features are input into a classifier of a support vector machine for prediction, and a prediction result is output, where the prediction result can predict the Ki-67 expression level of papillary thyroid carcinoma of the target object. The prediction method provided by the embodiment of the present invention can quickly and accurately predict the Ki-67 expression level of papillary thyroid carcinoma of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The exemplary embodiments of the present invention can be more fully understood by referring to the following drawings. The drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the present invention and do not constitute a limitation to the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0060] Figure 1 Schematic diagram of the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning according to the present invention;
[0061] Figure 2 Flowchart of the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided according to an exemplary embodiment of the present invention;
[0062] Figure 3 Flowchart of the image segmentation and preprocessing according to the present invention;
[0063] Figure 4 Basic architecture diagram of the SE-ResNet according to the present invention;
[0064] Figure 5 Structure diagram of the FPN according to the present invention;
[0065] Figure 6 Diagram of the CBAM module according to the present invention;
[0066] Figure 7 Structure diagram of the prediction device 700 for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided according to an exemplary embodiment of the present invention. Detailed implementation manners
[0067] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully communicated to those skilled in the art.
[0068] It should be noted that, unless otherwise specified, the technical terms or scientific terms used in the present invention should have the ordinary meanings understood by those skilled in the art to which the present invention belongs.
[0069] In addition, terms such as "first" and "second" are used to distinguish different objects, rather than to describe a specific order. Furthermore, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products, or devices.
[0070] Embodiments of the present invention provide a prediction method and apparatus, a computer-readable medium, and an electronic device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning, which will be described below with reference to the accompanying drawings.
[0071] As Figure 1 shown, it is a schematic diagram of the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning of the present invention.
[0072] As Figure 1 shown, the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by the embodiments of the present invention includes: Patient data collection: Obtain B-mode ultrasound images of patients with papillary thyroid carcinoma; Image segmentation and preprocessing: Segment the region of interest (ROI) of the tumor area in the B-ultrasound image, and use histogram equalization and the Canny operator to respectively characterize the texture and edge information of the lesion, and then perform channel fusion with the B-ultrasound image; Clinical information processing: Obtain the clinical information of the patient corresponding to the B-ultrasound image, perform data structuring processing and encoding to obtain effective clinical features; Deep learning radiomics feature extraction: Input the preprocessed image into a deep learning model, and at the same time use pre-trained weights for transfer learning and fine-tuning to obtain deep learning radiomics features; Classification model (classifier) training and prediction: Combine the deep learning radiomics features and clinical features to form combined features, and input the combined features into an SVM classifier for learning to obtain the prediction result of the Ki-67 expression level.
[0073] Please refer to Figure 2 , which shows a flowchart of the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by some embodiments of the present invention. As Figure 2 shown, the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning may include the following steps:
[0074] Step S201: Perform image preprocessing on any one of the multiple B-mode ultrasound images of the target object to obtain the corresponding preprocessed image, where the target object is a patient with papillary thyroid carcinoma.
[0075] In actual application scenarios, it is necessary to collect patient data: collect B-mode ultrasound images of patients with papillary thyroid carcinoma to ensure that the image quality meets the requirements of subsequent processing.
[0076] The specific steps for collecting patient data are as follows:
[0077] Step a1, Research object: In this retrospective analysis, 709 ultrasound images collected from 298 surgically confirmed PTC cases in Peking Union Medical College Hospital from September 2017 to October 2021 were included in this study. The images were divided into positive and negative image samples with the Ki-67 index value of 5% as the cut-off value (positive group: Ki-67 index > 5%, negative group: Ki-67 index ≤ 5%). The ratio of positive to negative image samples was approximately 6:4, and the dataset was divided into a training set and a test set according to 8:2.
[0078] Step a2, Inclusion criteria: PTC patients who received surgical treatment and had complete preoperative ultrasound images, and postoperative histopathological examinations included Ki-67 index values.
[0079] Step a3, Exclusion criteria: Patients without preoperative ultrasound images or with poor image quality, and those without Ki-67 index in postoperative pathology.
[0080] In one example, image preprocessing is performed on any one of multiple B-mode ultrasound images of a target object, including the following steps:
[0081] Randomly select any one of the multiple B-mode ultrasound images of the target object as the current image;
[0082] Perform the following preprocessing operations on the current image:
[0083] Perform preprocessing of image enhancement on the current image using histogram equalization and Canny operator to enhance the texture information and edge information of the corresponding image;
[0084] Perform image data enhancement processing on the current image using at least one of image data enhancement methods including random cropping, random horizontal flipping, and normalization processing to obtain the enhanced image;
[0085] Adopt a channel fusion method to combine the current image and the enhanced image for feature extraction;
[0086] Traverse the remaining images in the multiple B-mode ultrasound images of the target object except the current image, and perform preprocessing operations on any one of the remaining images.
[0087] In one example, the prediction method provided by the embodiments of the present invention may further include the following steps:
[0088] Obtain any one of multiple B-mode ultrasound images of a target object;
[0089] For any one of multiple B-mode ultrasound images of a target object, perform region of interest segmentation processing to extract the corresponding tumor region for image segmentation processing.
[0090] As Figure 3 shown, it is a flowchart of image segmentation and preprocessing of the present invention. As Figure 3 shown, perform region of interest (ROI) segmentation on the tumor region of the B-mode ultrasound image, and use histogram equalization and Canny operator to respectively represent the texture and edge information of the lesion, and then perform channel fusion with the B-mode ultrasound image.
[0091] Step S202: Construct a preset deep learning model.
[0092] In one example, the components of the preset deep learning model include: using the ResNet-50 model as the basic model architecture, integrating multiple target modules, and the multiple target modules include the SE module, the CBAM module, the FPN module, and the clinical attention module. The SE module is used to re-weight the channels of the feature map, the CBAM module is used to enhance the feature extraction performance, the FPN module is used to construct a multi-scale feature pyramid, and the clinical attention module is used to enhance the interaction process between the image features and the clinical features.
[0093] In one example, constructing a preset deep learning model includes the following steps:
[0094] Use the ResNet-50 model as the basic model architecture, and construct the corresponding SE-ResNet50 model through the SE module. Each residual block of the SE-ResNet50 model is replaced with the corresponding SE module version;
[0095] On the basis of the SE-ResNet50 model, fuse the CBAM module and the FPN module;
[0096] Define the way of data flowing through the network through forward propagation;
[0097] Enhance the interaction process between the image features and the clinical features through the clinical attention module.
[0098] In one example, using the ResNet-50 model as the basic model architecture and constructing the corresponding SE-ResNet50 model through the SE module includes the following steps:
[0099] Compress the spatial dimension in a global average pooling manner through the SE module;
[0100] Generate the weights for each channel through the SE module:
[0101] Generate the weights for each channel through a first fully-connected layer, a second fully-connected layer, and a Sigmoid activation function. The first fully-connected layer is used to reduce the number of channels, and the second fully-connected layer is used to restore to the original number of channels.
[0102] In an actual application scenario, the process of constructing the SE-ResNet50 model (as shown in Figure 4 which is the basic architecture diagram of the SE-ResNet of the present invention) is specifically as follows: A model based on ResNet-50 is constructed, where each residual block is replaced by a version of the SE module. Specifically: The SEBottleneck class is a residual block in SE-ResNet50, which contains three convolutional layers (1x1, 3x3, 1x1), and each convolutional layer is followed by a batch normalization layer (BatchNorm). After the last convolutional layer, an SE module is added to re-weight the channels of the feature map.
[0103] In one example, the SE module is implemented through the SELayer class. First, global average pooling is used to compress the spatial dimension, and then the weights for each channel are generated through two fully-connected layers (the first reduces the number of channels, and the second restores the original number of channels) and a Sigmoid activation function.
[0104] In one example, on the basis of the SE-ResNet50 model, the CBAM module and the FPN module are fused, including the following steps:
[0105] Obtain the SE-ResNet50 model and remove the original fully-connected layer of the SE-ResNet50 model;
[0106] Add an FPN module (as shown in Figure 5 which is the FPN structure diagram of the present invention). The FPN module extracts features from multiple different stages of the SE-ResNet50 model and constructs a multi-scale feature pyramid through upsampling and weighted summation;
[0107] For multiple different stages, add a number of CBAM modules that match the number of stages to enhance the channel and spatial attention of the feature map through the CBAM module, where each feature pyramid level corresponds to a CBAM module;
[0108] For the feature channels of each level of the feature pyramid, perform global average pooling and flattening processing, and perform classification processing through a linear classifier.
[0109] In an actual application scenario, the process of integrating the CBAM module and the FPN module is specifically as follows: First, the pre-trained SE-ResNet50 model constructed through step 51 is used, and the original fully connected layer is removed. Then, an FPN module is added, which extracts features from different stages (c1, c2, c3, c4) of the SE-ResNet50 and constructs a multi-scale feature pyramid through upsampling and weighted summation. Next, four CBAM modules (one for each feature pyramid level) are added, which enhance the channel and spatial attention of the feature map. Finally, the features of each level of the feature pyramid are globally average pooled and flattened, and then classified through a linear layer.
[0110] In an actual application scenario, the process of forward propagation is specifically as follows:
[0111] Forward propagation: Defines how data flows through the network, including extracting features from different stages (c1, c2, c3, c4) through the base model of the SE-ResNet50, feeding them into the FPN module to generate a multi-scale feature pyramid (p1, p2, p3, p4), and each level of the feature pyramid is processed through a CBAM module (as Figure 6 shown, it is the CBAM module diagram of the present invention), enhancing the channel and spatial attention of the features, and finally globally average pooling and flattening them and then feeding them into a linear classifier for classification.
[0112] In one example, the process of enhancing the interaction between image features and clinical features through the clinical attention module includes the following steps:
[0113] Mapping the image features to the clinical feature space and mapping the clinical features to the image feature space to implement the processing process of feature space mapping;
[0114] Calculating the attention weights through the ReLU activation function and the Sigmoid activation function;
[0115] Based on the attention weights, performing weighted processing on the image features and the clinical features to enhance the first representation ability of the image features and the second representation ability of the clinical features.
[0116] In one example, the prediction method provided by the embodiments of the present invention may further include the following steps:
[0117] Setting the hyperparameters of the preset deep learning model.
[0118] In one example, setting the hyperparameters of the preset deep learning model includes the following steps:
[0119] Setting the maximum number of iterations to 600;
[0120] Set the learning rate to 2e-5;
[0121] Set the optimizer to the Adam optimizer;
[0122] Set the decay weight to an L2 regularization term of 1e-3 to control the magnitude of the parameters by the decay weight to prevent overfitting;
[0123] Set the updated batch size to 32.
[0124] In one example, the preset deep learning model adopts a cross-entropy loss function, and the cross-entropy loss function is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label.
[0125] In one example, the preset deep learning model adopts a learning rate decay strategy; and,
[0126] The preset deep learning model adopts a ReduceLROnPlateau scheduler. Through the ReduceLROnPlateau scheduler, if the loss does not improve within a preset number of consecutive cycles, the corresponding learning rate will decay to a preset value.
[0127] In an actual application scenario, the preset deep learning model adopts a ReduceLROnPlateau scheduler. When the loss does not improve within 10 consecutive cycles, the learning rate will decay to half of the original.
[0128] The above preset value can be adjusted according to the requirements of different application scenarios and is not specifically limited here.
[0129] Step S203: Input the preprocessed image into the preset deep learning model for feature extraction to obtain target radiomics features. The target radiological features are radiomics features that can reflect the Ki-67 expression level of papillary thyroid carcinoma of the target object.
[0130] Step S204: Obtain the processed clinical features of the target object and obtain the target radiomics features.
[0131] In one example, obtaining the processed clinical features of the target object includes the following steps:
[0132] Obtain any one of the multiple B-mode ultrasound images of the target object;
[0133] For any one B-mode ultrasound image, obtain the corresponding clinical information and perform data structuring processing and coding processing on the clinical information to obtain the processed clinical features of the target object. The processed clinical features of the target object are valid clinical features.
[0134] In actual application scenarios, before obtaining the processed clinical characteristics of the target object, the clinical information needs to be processed: collect the clinical information of the patient corresponding to the B-ultrasound image, including age, gender, height, weight, preoperative thyroid function indicators (FT3, FT4, TSH, Tg-Ab, TPO-Ab), whether there is Hashimoto's thyroiditis or nodular goiter, and perform structured processing and encoding of the data to facilitate combination with deep learning imaging genomics features.
[0135] Step S205: performing feature combination processing on the processed clinical features and target radiomics features to obtain corresponding joint features.
[0136] Step S206: input the joint features into the classifier of the support vector machine for prediction, and output the prediction result, which can predict the Ki-67 expression level of papillary thyroid carcinoma of the target object.
[0137] In one example, the prediction method provided by the embodiment of the present invention further includes the following steps:
[0138] Acquire a plurality of B-mode ultrasound images of a plurality of different objects to be tested;
[0139] Based on a plurality of B-mode ultrasound images of a plurality of different objects to be tested, a plurality of sample data for training a classifier capable of supporting a vector machine are collected;
[0140] According to a first preset ratio, the plurality of sample data are divided into positive image sample data and negative image sample data, the Ki-67 index corresponding to any positive image sample data is greater than the first preset value, and the Ki-67 index corresponding to any negative image sample data is less than or equal to the first preset value;
[0141] According to a second preset ratio, the plurality of sample data are divided into training sample data for training the classifier and testing sample data for testing the classifier.
[0142] In actual application scenarios, the classification model (classifier) training and prediction process is specifically described as follows: the extracted deep learning radiomics features and the processed clinical features are combined into joint features, which are input into the support vector machine (SVM) classifier for training and learning, and finally the prediction results of Ki-67 expression levels are obtained.
[0143] In the prediction method provided in the embodiment of the present invention, a clinical attention module ClinicalAttentionModule is designed to enhance the interaction between image features and clinical features through the attention mechanism. The specific implementation is as follows:
[0144] Step b1, Feature space mapping: First, map the image features to the clinical feature space and map the clinical features to the image feature space.
[0145] Step b2, Attention weight calculation: Use the ReLU activation function and the Sigmoid activation function to calculate the attention weights.
[0146] Step b3, Feature enhancement: Use the calculated attention weights to weight the image features and clinical features, thereby enhancing their representation ability.
[0147] In one example, the SVM classifier is a random forest classifier after training and validation.
[0148] In one example, the training of the SVM classifier uses GridSearchCV for parameter search to find the best parameter combination, including C, kernel, gamma, and degree. Here, 5-fold cross-validation and accuracy are used as the scoring criteria to adjust the weight parameters of the SVM classifier to obtain the trained and validated SVM classifier.
[0149] In one example, the ratio of the training set to the test set used in training and validation is 8:2, and the method of 5-fold cross-validation is used for training and validation.
[0150] In one example, when testing the SVM classifier, accuracy (Accuracy, ACC), sensitivity (Sensitivity, SEN), specificity (Specificity, SPE), and AUC are used as evaluation criteria.
[0151] In one example, when the random forest classifier is trained and validated, the B-ultrasound images and original clinical information data used are the preoperative data and postoperative pathological data of patients with surgically confirmed papillary thyroid carcinoma.
[0152] In one example, using the same set of B-ultrasound images and original clinical information data, the combined features obtained based on the proposed deep learning radiomics model are input into the same SVM classifier for classification, and the following results are obtained:
[0153] In the Ki-67 risk level classification task, the accuracy of the prediction results obtained by the prediction method provided by the embodiment of the present invention is 75.4%, the sensitivity is 58.9%, the specificity is 86.1%, and the AUC is 72.5%.
[0154] The prediction method provided by the embodiment of the present invention can quickly and accurately predict the Ki-67 expression level of papillary thyroid carcinoma in the target object.
[0155] In the above embodiments, a prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning is provided. Correspondingly, the present invention also provides a prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning. The prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by the embodiments of the present invention can implement the above prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning, and the prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning can be implemented in a software, hardware, or a combination of software and hardware manner. For example, the prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning can include integrated or separate functional modules or units to execute the corresponding steps in the above various methods.
[0156] Please refer to Figure 7 , which shows a schematic diagram of a prediction device for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by some embodiments of the present invention. Since the device embodiments are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, please refer to the partial description of the method embodiments. The device embodiments described below are only illustrative.
[0157] As Figure 7 shown, the prediction device 700 for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning may include:
[0158] A preprocessing module 701, configured to perform image preprocessing on any one of a plurality of B-mode ultrasound images of a target object to obtain a corresponding preprocessed image, where the target object is a patient with papillary thyroid carcinoma;
[0159] A model construction module 702, configured to construct a preset deep learning model;
[0160] A feature extraction module 703, configured to input the preprocessed image into the preset deep learning model for feature extraction to obtain target radiomics features, where the target radiological features are radiomics features that can reflect the Ki-67 expression level of the papillary thyroid carcinoma of the target object;
[0161] An acquisition module 704, configured to acquire the processed clinical features of the target object and acquire the target radiomics features;
[0162] A feature combination module 705, configured to perform feature combination processing on the processed clinical features and the target radiomics features to obtain corresponding combined features;
[0163] A prediction module 706 is configured to input the combined features into a classifier of a support vector machine for prediction and output a prediction result, which can predict the Ki-67 expression level of papillary thyroid carcinoma of a target object.
[0164] In some embodiments of the embodiments of the present invention, the components of the preset deep learning model include: using the ResNet-50 model as the basic model architecture, integrating multiple target modules, and the multiple target modules include an SE module, a CBAM module, an FPN module, and a clinical attention module. The SE module is used to re-weight the channels of the feature map, the CBAM module is used to enhance the feature extraction performance, the FPN module is used to construct a multi-scale feature pyramid, and the clinical attention module is used to enhance the interaction process between the image features and the clinical features.
[0165] In some embodiments of the embodiments of the present invention, the model construction module 702 is specifically configured to:
[0166] Using the ResNet-50 model as the basic model architecture, constructing a corresponding SE-ResNet50 model through the SE module, and each residual block of the SE-ResNet50 model is replaced with a corresponding SE module version;
[0167] On the basis of the SE-ResNet50 model, integrating the CBAM module and the FPN module;
[0168] Defining the flow mode of data through the network through forward propagation;
[0169] Enhancing the interaction process between the image features and the clinical features through the clinical attention module.
[0170] In some embodiments of the embodiments of the present invention, the model construction module 702 is specifically configured to:
[0171] Mapping the image features to the clinical feature space and mapping the clinical features to the image feature space to implement the processing process of feature space mapping;
[0172] Calculating the attention weights through the ReLU activation function and the Sigmoid activation function;
[0173] Based on the attention weights, performing weighted processing on the image features and the clinical features to enhance the first representation ability of the image features and the second representation ability of the clinical features.
[0174] In some embodiments of the embodiments of the present invention, the model construction module 702 is specifically configured to:
[0175] Obtain the SE-ResNet50 model and remove the original fully connected layer of the SE-ResNet50 model;
[0176] Add an FPN module. The FPN module extracts features from multiple different stages of the SE-ResNet50 model and constructs a multi-scale feature pyramid through upsampling and weighted summation;
[0177] For multiple different stages, add multiple CBAM modules that match the number of stages to enhance the channel and spatial attention of the feature map through the CBAM modules. Among them, each level of the feature pyramid corresponds to one CBAM module;
[0178] For the feature channels of each level of the feature pyramid, perform global average pooling and flattening, and perform classification through a linear classifier.
[0179] In some embodiments of the embodiments of the present invention, the model construction module 702 is specifically configured to:
[0180] Compress the spatial dimension in a global average pooling manner through the SE module;
[0181] Generate the weight of each channel through the SE module:
[0182] Generate the weight of each channel through a first fully connected layer, a second fully connected layer, and a Sigmoid activation function. The first fully connected layer is used to reduce the number of channels, and the second fully connected layer is used to restore to the original number of channels.
[0183] In some embodiments of the embodiments of the present invention, the prediction device 700 for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by the embodiments of the present invention may further include:
[0184] A setting module (not shown in Figure 7 ) for
[0185] Set the hyperparameters of the preset deep learning model;
[0186] The setting module is specifically configured to: set the maximum number of iterations to 600; set the learning rate to 2e-5; set the optimizer to the Adam optimizer; set the decay weight to the L2 regularization term of 1e-3 to control the amplitude of the parameters through the decay weight to prevent overfitting; and set the updated batch size to 32.
[0187] In some embodiments of the embodiments of the present invention, the preset deep learning model adopts a cross-entropy loss function, and the cross-entropy loss function is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label; alternatively, the preset deep learning model adopts a learning rate decay strategy; and, the preset deep learning model adopts a ReduceLROnPlateau scheduler, so that through the ReduceLROnPlateau scheduler, if the loss has not improved within a preset number of consecutive cycles, the corresponding learning rate will decay to a preset value.
[0188] In some embodiments of the embodiments of the present invention, the preprocessing module 701 is specifically configured to:
[0189] Randomly select any one B-mode ultrasound image from multiple B-mode ultrasound images of the target object as the current image;
[0190] Perform the following preprocessing operations on the current image:
[0191] Perform preprocessing of image enhancement on the current image using histogram equalization and Canny operator to enhance the texture information and edge information of the corresponding image;
[0192] Perform image data enhancement processing on the current image using at least one image data enhancement method including random cropping, random horizontal flipping, and normalization processing to obtain an enhanced image;
[0193] Adopt a channel fusion method to combine the current image and the enhanced image for feature extraction;
[0194] Traverse the remaining images in the multiple B-mode ultrasound images of the target object except the current image, and perform preprocessing operations on any one of the remaining images.
[0195] In some embodiments of the embodiments of the present invention, in some embodiments of the embodiments of the present invention, the prediction device 700 for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by the embodiments of the present invention and the prediction method for the Ki-67 expression level of papillary thyroid carcinoma based on deep learning provided by the foregoing embodiments of the present invention are based on the same inventive concept and have the same beneficial effects.
[0196] According to an embodiment of another aspect, there is also provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed in a computer, the computer is made to execute the method described in combination with Figure 1 described above.
[0197] According to an embodiment of still another aspect, an electronic device is further provided, including a memory and a processor. An executable code is stored in the memory. When the processor executes the executable code, the method described in conjunction with Figure 1 is implemented.
[0198] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0199] The specific embodiments described above have further elaborated on the purpose, technical solutions, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for predicting Ki-67 expression level in papillary thyroid carcinoma based on deep learning, characterized in that: The method comprises: Performing image preprocessing on any one of a plurality of B-mode ultrasound images of a target object to obtain a corresponding preprocessed image, wherein the target object is a patient with papillary thyroid cancer; Build preset deep learning models; Inputting the preprocessed image into the preset deep learning model for feature extraction to obtain target imaging features, wherein the target imaging features are imaging features that can reflect the Ki-67 expression level of papillary thyroid carcinoma of the target object; Obtaining the treated clinical features of the target object, and obtaining the target radiomic features; Performing feature combination processing on the processed clinical features and the target radiomics features to obtain corresponding joint features; The joint feature is input into a classifier of a support vector machine for prediction, and a prediction result is output, wherein the prediction result can predict the Ki-67 expression level of thyroid papillary carcinoma of the target object.
2. The prediction method according to claim 1, characterized in that: The components of the preset deep learning model include: using the ResNet-50 model as the basic model architecture, integrating multiple target modules, the multiple target modules include an SE module, a CBAM module, an FPN module and a clinical attention module, the SE module is used to reweight the channels of the feature map, the CBAM module is used to enhance the feature extraction performance, the FPN module is used to construct a multi-scale feature pyramid, and the clinical attention module is used to enhance the interaction process between image features and clinical features.
3. The prediction method according to claim 2, characterized in that: The step of constructing the preset deep learning model includes: The ResNet-50 model is used as the basic model architecture, and the corresponding SE-ResNet50 model is constructed through the SE module, and each residual block of the SE-ResNet50 model is replaced by the corresponding SE module version; Based on the SE-ResNet50 model, the CBAM module and the FPN module are integrated; Define how data flows through the network through forward propagation; The interaction process between image features and clinical features is enhanced by the clinical attention module.
4. The prediction method according to claim 3, characterized in that: The process of enhancing the interaction between image features and clinical features by the clinical attention module comprises: Mapping image features to clinical feature space, and mapping clinical features to image feature space, to achieve feature space mapping process; The attention weight is calculated by ReLU activation function and Sigmoid activation function; Based on the attention weights, the image features and the clinical features are weighted to enhance the first representation capability of the image features and the second representation capability of the clinical features.
5. The prediction method according to claim 3, characterized in that: Based on the SE-ResNet50 model, the CBAM module and the FPN module are integrated, including: Obtain the SE-ResNet50 model, and remove the original fully connected layer of the SE-ResNet50 model; Adding an FPN module, which extracts features from multiple different stages of the SE-ResNet50 model and constructs a multi-scale feature pyramid by upsampling and weighted summation; For the multiple different stages, add a plurality of the CBAM modules matching the number of stages, so as to enhance the channel and spatial attention of the feature map through the CBAM modules, wherein each feature pyramid level corresponds to one CBAM module; For the feature channels at each level of the feature pyramid, global average pooling and flattening are performed, and classification is performed through a linear classifier.
6. The prediction method according to claim 3, characterized in that: The ResNet-50 model is used as the basic model architecture, and the corresponding SE-ResNet50 model is constructed through the SE module, including: Through the SE module, the spatial dimension is compressed in a global average pooling manner; The weight of each channel is generated by the SE module: The weight of each channel is generated through a first fully connected layer, a second fully connected layer and a Sigmoid activation function, wherein the first fully connected layer is used to reduce the number of channels, and the second fully connected layer is used to restore the original number of channels.
7. The prediction method according to claim 1, characterized in that: The method further comprises: Set hyperparameters of a preset deep learning model; The hyperparameters of the preset deep learning model are as follows: Set the maximum number of iterations to 600; Set the learning rate to 2e-5; Set the optimizer to Adam optimizer; The decay weight is set to an L2 regularization term of 1e-3 to control the magnitude of the parameter by the decay weight to prevent overfitting; Set the batch size for updates to 32.
8. The prediction method according to claim 1, characterized in that: The preset deep learning model adopts a cross entropy loss function, which is used to measure the difference between the probability distribution predicted by the model and the probability distribution of the true label; or, The preset deep learning model adopts a learning rate decay strategy; and, The preset deep learning model adopts the ReduceLROnPlateau scheduler, so that if the loss does not improve within a preset number of consecutive cycles through the ReduceLROnPlateau scheduler, the corresponding learning rate will decay to a preset value.
9. The prediction method according to claim 1, characterized in that: The performing image preprocessing on any one of the multiple B-mode ultrasound images of the target object comprises: Randomly select any one B-mode ultrasound image from a plurality of B-mode ultrasound images of the target object as the current image; Perform the following preprocessing operations on the current image: The current image is preprocessed by using histogram equalization and Canny operator to enhance the texture information and edge information of the corresponding image; Performing image data enhancement processing on the current image by using at least one image data enhancement method including random cropping, random horizontal flipping, and normalization processing to obtain an enhanced image; The channel fusion method is used to combine the current image and the enhanced image for feature extraction; The remaining images except the current image in the multiple B-mode ultrasound images of the target object are traversed, and a preprocessing operation is performed on any one of the remaining images.
10. A prediction device for Ki-67 expression level of papillary thyroid carcinoma based on deep learning, characterized in that: The device comprises: A preprocessing module, used for performing image preprocessing on any one of a plurality of B-mode ultrasound images of a target object to obtain a corresponding preprocessed image, wherein the target object is a patient with papillary thyroid cancer; Model building module, used to build preset deep learning models; A feature extraction module, used for inputting the preprocessed image into the preset deep learning model for feature extraction to obtain a target imaging feature, wherein the target imaging feature is an imaging feature that can reflect the Ki-67 expression level of papillary thyroid carcinoma of the target object; An acquisition module, used to acquire the processed clinical features of the target object and to acquire the target radiomic features; A feature combination module, used for performing feature combination processing on the processed clinical features and the target imaging features to obtain corresponding joint features; The prediction module is used to input the joint features into the classifier of the support vector machine for prediction and output the prediction result, wherein the prediction result can predict the Ki-67 expression level of thyroid papillary carcinoma of the target object.
Citation Information
Patent Citations
Pulmonary nodule benign and malignant prediction method and device
CN111915596A
Soft tissue opto-acoustic / ultrasonic multi-modal image fusion method based on deep learning
CN118212495A
Multi-modal emotion recognition method based on state space model cross-modal interaction
CN119128578A
Emotion analysis method and system based on multi-modal fusion
CN119272224A