Thyroid tumor diagnosis method and system based on ultrasonic and cytological image conjoint analysis
By combining B-ultrasound images and cytological images with the Transformer architecture for multimodal fusion analysis, the problem of limited diagnostic performance of a single modality was solved, and the accuracy and efficiency of thyroid tumor diagnosis were improved.
Patent Information
- Application Number
- CN202510726372.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies rely on single-modality image analysis in thyroid tumor diagnosis, resulting in limited diagnostic performance, especially poor performance in cases of unclear boundaries or heterogeneous lesions, and lack of multimodal information fusion.
A joint analysis method of ultrasound and cytology images is used to convert B-ultrasound images and cytology images into image feature vectors and structured feature vectors, which are then fused through a multimodal prediction network with a Transformer architecture to generate a fused Token sequence for diagnosis.
It significantly improves the accuracy of thyroid tumor diagnosis, especially the detection sensitivity of microcancer and borderline cases by 5% to 10%, reduces manual operation costs, and supports real-time bedside analysis.
Smart Images

Figure CN120612533A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of thyroid tumor diagnosis, and in particular to a thyroid tumor diagnosis method and system for combined analysis of ultrasound and cytological images. Background Art
[0002] Preoperative diagnosis of thyroid tumors is crucial for developing treatment strategies. Currently, clinical diagnosis relies primarily on ultrasound imaging and cytological puncture biopsy. With the development of deep learning technology, intelligent assisted diagnosis based on medical images has become a research hotspot, but existing technologies still have significant limitations.
[0003] Prior art Chinese patent application CN201810318306.X discloses a thyroid tumor ultrasound image recognition method and device thereof, the method comprising: selecting a tumor area in a thyroid tumor ultrasound image and enlarging a certain edge range before cutting, performing benign and malignant labeling, and forming the cut images into a training set; using the training set to train a selected convolutional neural network to form a thyroid tumor ultrasound image recognition model; obtaining a thyroid tumor ultrasound image to be identified, selecting a tumor area and enlarging a certain edge range, and then using the thyroid tumor ultrasound image recognition model to perform benign and malignant identification.
[0004] This existing technology uses convolutional neural networks to analyze tumor regions in ultrasound images, but it only utilizes the structural features of ultrasound and lacks microscopic information about cell morphology. It also relies on manual delineation of tumor regions and amplification of margins, without involving quantitative analysis at the cellular level. Furthermore, it inadequately processes noise in ultrasound images, impacting the quality of feature extraction.
[0005] There is also prior art Chinese patent application CN201810318298.9 that discloses a network construction method and system for thyroid tumor cytology smear image classification. The system uses a reinforcement learning method to find the existing convolutional neural network that is most suitable for thyroid tumor cytology smear image classification. The specific process of the reinforcement learning method is: first, a convolutional neural network is generated using a recurrent neural network; then, the convolutional neural network is trained using a thyroid tumor cytology smear image training set; then, the accuracy of the trained convolutional neural network is verified using a thyroid tumor cytology smear image verification set, and an accuracy threshold is set to determine whether its accuracy is higher than the threshold; finally, the convolutional neural network with the highest accuracy is retrained as a preliminary convolutional neural network, thereby achieving the purpose of constructing a high-accuracy convolutional neural network to assist doctors in diagnosing thyroid tumors and improving the diagnostic accuracy.
[0006] This existing technology proposes to classify cytology smear images based on reinforcement learning to optimize convolutional networks, but it is only based on single-modality cytology data and cannot integrate the overall distribution characteristics of tumors in ultrasound images.
[0007] It can be seen that most existing methods rely on a single modality, such as B-ultrasound images in CN201810318306.X or cytology images in CN201810318298.9. Due to the incomplete information of the image itself, the diagnostic performance of the model has an upper limit, especially in the case of unclear boundaries or heterogeneous lesions. Summary of the Invention
[0008] The purpose of the present invention is to provide a thyroid tumor diagnosis method and system for joint analysis of ultrasound and cytological images, which partially solves or alleviates the above-mentioned shortcomings in the prior art and can perform thyroid tumor diagnosis based on multimodal fusion of ultrasound and cytological images, thereby improving the accuracy of diagnosis.
[0009] In order to solve the above-mentioned technical problems, the present invention specifically adopts the following technical solutions:
[0010] A first aspect of the present invention is to provide a method for diagnosing thyroid tumors by combining analysis of ultrasound and cytological images, comprising:
[0011] Convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US ;
[0012] The structured feature vector F is obtained by extracting the structured feature of the thyroid cytology image of the patient to be diagnosed. CYTO ;
[0013] For the image feature vector F US And the structured feature vector F CYTO Convert to fusion token sequence;
[0014] It is used to input the fused token sequence into the multimodal prediction network classification prediction based on the Transformer architecture, so as to obtain the probability that the thyroid tumor of the patient to be diagnosed is malignant.
[0015] Furthermore, the thyroid ultrasound image of the patient to be diagnosed is converted into an image feature vector F US The steps include:
[0016] Preprocessing of thyroid B-ultrasound images of patients to be diagnosed;
[0017] The EfficientNet-B0 model is trained using transfer learning;
[0018] The preprocessed image is input into the pre-trained EfficientNet-B0 model for image feature extraction to obtain the image feature vector F US .
[0019] Furthermore, the step of preprocessing the thyroid B-ultrasound image of the patient to be diagnosed includes:
[0020] Crop the image edges according to the preset ratio;
[0021] Perform white balance processing and grayscale conversion;
[0022] The image is converted into a single-channel grayscale image through image binarization and morphological dilation operations;
[0023] The thyroid diagnostic area is located in the image through contour detection, and the thyroid diagnostic area image is obtained by cropping;
[0024] Scale the image to a preset size and normalize it.
[0025] Furthermore, before training the EfficientNet-B0 model, image enhancement is performed on the dataset used for training;
[0026] During training, the benign / malignant labels are used as supervisory signals to train the model. The training parameters include: Adam optimizer, initial learning rate of 1e-4, batch size of 32, binary cross entropy loss function, 100 training rounds, early stopping mechanism and CosineAnnealingLR learning rate adjustment strategy.
[0027] The structured feature vector F is obtained by extracting the structured feature of the thyroid cytology image of the patient to be diagnosed. CYTO The steps include:
[0028] The YOLOv5 object detection model was trained on cytology WSI images.
[0029] Use the trained YOLOv5 target detection model to detect and count the multiple types of cells in the thyroid cytology WSI image, extract the structural information and generate the structural feature vector F CYTO .
[0030] Furthermore, the step of training the YOLOv5 target detection model includes:
[0031] A sliding window strategy is used to crop the WSI image to generate image patches of fixed size;
[0032] Using the cell category and location annotation data provided by pathology experts, we constructed a cell detection training dataset.
[0033] The YOLOv5 model was selected as the target detection network structure, and the model was adapted and trained based on the characteristics of medical images. The adaptation included setting a smaller anchor matching threshold, a lower initial learning rate, and a reasonable data augmentation strategy (such as rotation and hue adjustment).
[0034] With the positioning and classification of cell targets as the training goal, the detection and recognition tasks of multiple categories including nuclear division cells, atypical cell nuclei, papillary carcinoma characteristic cells, follicular carcinoma characteristic cells and other background or atypical cells are achieved.
[0035] Furthermore, the structured feature vector F is generated CYTO The steps include:
[0036] The trained YOLOv5 target detection model is used to detect and classify multiple cell types in thyroid cytology WSI images. Statistical analysis is performed based on the detection results to extract structured statistical features of the images, including but not limited to: cell type statistics, average detection score statistics, average confidence statistics, median detection statistics, and median position confidence statistics.
[0037] The above statistical features are aggregated per patient and processed by one-hot encoding and other methods to finally generate a fixed-dimensional structured feature vector F CYTO .
[0038] Furthermore, the B-ultrasound image feature vector F US and cytological structured feature vector F CYTO The steps for converting to a fusion token sequence include:
[0039] The image feature vector F US The feature map represented is divided into n spatial regions through adaptive average pooling, obtaining n image feature vector units; each unit is mapped to a p-dimensional vector space through linear transformation to form n image tokens; the corresponding modality category code and position code are added to each image token;
[0040] The structured feature vector F CYTO Divide into m structured feature vector units according to cell type, and map each unit to p-dimensional vector space through linear transformation to form m structured tokens; modal category code and position code are also added to each structured token;
[0041] The n image tokens are concatenated with the m structured tokens to form a fused token sequence with a length of n+m and a dimension of p.
[0042] Furthermore, the step of inputting the fused Token sequence into a multimodal prediction network based on the Transformer architecture for classification prediction includes:
[0043] The fused token sequence formed by concatenating the ultrasound image token and the cytology image token is input into a multimodal encoder composed of a stack of multi-layer Transformer encoders. The Transformer encoder includes a multi-head self-attention mechanism, a feedforward neural network sublayer, a residual connection, and a normalization mechanism to model the contextual dependencies between cross-modal features and enhance the fusion representation capability.
[0044] The Token sequence output by the Transformer encoding is subjected to an average pooling operation to generate a fused feature vector;
[0045] The fused feature vector is input into a classifier comprising a two-layer fully connected network, which sequentially completes feature compression and malignancy probability prediction, and finally outputs the predicted probability of the patient having a malignant thyroid tumor.
[0046] Furthermore, the classifier uses the AdamW optimizer with an initial learning rate of 1e-4, combined with a cosine annealing learning rate scheduler; uses a batch size of 16, a total number of training rounds of 100 rounds, and an early stopping mechanism to avoid overfitting; the model training adopts a five-fold cross-validation method, and the evaluation indicators include accuracy, AUC, and F1 score.
[0047] The present invention also provides a thyroid tumor diagnosis system for combined analysis of ultrasound and cytological images, which is characterized by comprising:
[0048] The B-ultrasound image feature vector extraction module is used to convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US ;
[0049] The cytology image structured feature extraction module is used to extract the structured features of the thyroid cytology image of the patient to be diagnosed and obtain the structured feature vector F CYTO ;
[0050] Feature vector fusion module, used to fusion the image feature vector F US And the structured feature vector F CYTO Convert to fusion token sequence;
[0051] The prediction module is used to input the fused token sequence into the multimodal prediction network based on the Transformer architecture for classification prediction, thereby obtaining the probability that the thyroid tumor of the patient to be diagnosed is malignant.
[0052] Beneficial effects:
[0053] 1. Efficiency of preprocessing and feature extraction.
[0054] Through edge cropping, binarization, morphological expansion and contour detection, irrelevant areas such as device noise and black edges in ultrasound images are effectively removed, the boundaries of thyroid tissue are accurately located, and a high-purity diagnostic area mask is generated to ensure that subsequent analysis focuses on effective information.
[0055] Mask cropping and standardized unified image format adapt to the input requirements of deep learning models and improve the stability of feature extraction.
[0056] The EfficientNet-B0 pre-trained model is used to extract the image feature vector (F US ), using the common visual features (such as edges and textures) learned from natural images, combined with thyroid data fine-tuning, it significantly improves feature expression capabilities and reduces dependence on large-scale medical data.
[0057] By performing an adaptive average pooling operation on the feature map, it is divided into multiple spatial regions and regional features are extracted; each regional feature is mapped to a unified dimensional space through a linear transformation to form a token representation, retaining deep semantic information related to diagnosis, such as nodule morphology and echo pattern.
[0058] 2. Comprehensiveness of multimodal feature fusion.
[0059] Image tokens are derived from unstructured data such as ultrasound images, capturing morphological details of thyroid nodules (such as boundary characteristics and calcification distribution). Structured tokens are derived from cellular statistical features detected in cell images (such as the density of different cell types and the distribution of atypia), providing biological indicators at the microscopic level of pathology. The complementary fusion of these two modalities creates a comprehensive, integrated diagnostic expression spanning "macrostructure and microscopic pathology."
[0060] Through linear transformation, different modal features are uniformly mapped to the same dimensional space. Combined with explicitly injected modal category encoding and position encoding, the problems of semantic distribution differences and structural heterogeneity are effectively solved, and the model's modeling ability and interpretability of cross-modal relationships are improved.
[0061] The multi-layer, multi-head self-attention module based on the Transformer architecture integrates token sequence input to fully model long-range dependencies between and within modalities, simulating the logical reasoning process of clinicians integrating multi-source information during the diagnosis process.
[0062] 3. The advancement of the model architecture and its clinical adaptability.
[0063] Two layers of stacked Transformer encoders are used to implement a hierarchical processing architecture of "local feature extraction-global semantic fusion": the bottom-level encoder is used to model the local relationship between tokens within the modality, and the high-level encoder realizes deep interactive integration of information between modalities, enhancing the robustness and semantic richness of diagnostic expression.
[0064] The average pooling operation is applied to the Token sequence output by the Transformer code to obtain a fixed-dimensional fusion feature vector, which suppresses local abnormal signals, enhances the overall trend expression, and reduces computing resource overhead, making it suitable for clinical real-time diagnosis scenarios.
[0065] The subsequent two-layer fully connected network structure is concise and efficient: the first layer implements feature compression and activation, the second layer outputs a single probability value, and finally the Sigmoid activation function is used to represent the predicted probability of thyroid tumor malignancy. The output results are intuitive and clear, facilitating clinical risk assessment and auxiliary decision-making.
[0066] 4. Technical performance and clinical application value.
[0067] The multimodal fusion model significantly improves diagnostic performance by integrating complementary information. Experiments have shown that its classification accuracy is 5% to 10% higher than that of a single-modality model, with a particularly significant increase in sensitivity for detecting microcancers and borderline cases.
[0068] Automated preprocessing and feature extraction reduce manual operation costs, improve diagnostic efficiency, and support real-time bedside analysis; the modular design facilitates integration into existing medical systems and is compatible with data from multiple departments such as ultrasound and pathology.
[0069] The technical architecture is versatile and can be extended to other medical imaging or multi-omics data fusion. It is suitable for the diagnosis of thyroid tumors and has broad clinical application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the embodiments or the description of the prior art. In all drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the various elements or parts are not necessarily drawn according to the actual scale. Obviously, the drawings described below are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without inventive work.
[0071] Figure 1 This is a flow chart of embodiment 1 of the present invention.
[0072] Figure 2 This is a comparison chart of the accuracy of the present invention and the existing prediction method. DETAILED DESCRIPTION
[0073] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0074] Herein, suffixes such as "module," "component," or "unit" used to represent elements are only used to facilitate description of the present invention and have no specific meaning. Therefore, "module," "component," or "unit" may be used interchangeably.
[0075] As used herein, terms such as "upper," "lower," "inner," "outer," "front," "back," "one end," and "the other end" indicate positions or locations based on those shown in the accompanying drawings. These terms are intended solely to facilitate and simplify the description of the present invention and are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0076] As used herein, unless otherwise expressly specified or limited, the terms "installed," "provided with," and "connected" should be understood broadly. For example, "connected" may refer to a fixed connection, a detachable connection, or an integral connection; it may refer to a mechanical connection, a direct connection, an indirect connection via an intermediate medium, or internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention on a case-by-case basis.
[0077] As used herein, "and / or" includes any and all combinations of one or more of the associated listed items.
[0078] Herein, "plurality" means two or more than two, ie, it includes two, three, four, five, etc.
[0079] Example 1:
[0080] like Figure 1 As shown, this embodiment provides a thyroid tumor diagnosis method based on combined analysis of ultrasound and cytological images, including:
[0081] S1 converts the thyroid ultrasound image of the patient to be diagnosed into an image feature vector F US .
[0082] The purpose of this step is to convert medical images into machine-processable feature representations using computer vision technology. The specific steps include:
[0083] S11 pre-processes the thyroid ultrasound images of the patient to be diagnosed, specifically including:
[0084] S111 crops the image edges according to a preset ratio.
[0085] In this embodiment, a preprocessing operation is performed by cutting off a certain percentage or fixed width of edge regions from all sides of the image to remove irrelevant pixels and focus on the core content. In the field of medical imaging, edge regions often contain non-diagnostic information (i.e., "black edges") such as device identifiers, timestamps, and scale lines. This content is not only unhelpful for disease diagnosis but may also interfere with subsequent image analysis algorithms.
[0086] Black edges on thyroid B-ultrasound images may contain irrelevant information such as the ultrasound device model, patient ID, and procedure time, or even edge artifacts (such as shadows caused by probe pressure). Cropping can eliminate these noise artifacts that could mislead the model. The width of black edges on images generated by different ultrasound devices may vary. A uniform 10% edge cropping can reduce device variability, standardize the image preprocessing process, and improve the generalization of subsequent deep learning models.
[0087] S112 performs white balance processing and grayscale conversion.
[0088] White balancing corrects image color deviations, ensuring that "white" objects appear true white in the image and eliminating color casts caused by variations in light source color temperature. Differences in imaging sensors or probes across ultrasound devices can cause images to appear cooler or warmer, affecting the interpretation of tissue echogenicity. White balancing is essential for consistent color styling.
[0089] In this embodiment, the grayscale world hypothesis or an algorithm based on RGB dynamic range equalization can be used to perform white balance correction on the image to solve the problem of inconsistent color temperature caused by device differences and improve the style consistency between images.
[0090] Grayscale conversion converts RGB color images into single-channel grayscale images, removing color redundancy while retaining brightness information. This makes it easier to capture key structural features such as nodule boundaries and calcifications in subsequent processing.
[0091] S113 converts the image into a single-channel grayscale image through image binarization and morphological dilation operations.
[0092] Image binarization converts a grayscale image into an image containing only two pixel values: "black" (0) and "white" (255). This separation is achieved by setting a threshold. In this example, the Otsu method is used to automatically determine the optimal threshold by maximizing the between-class variance. This method is suitable for images with a bimodal histogram.
[0093] Morphological dilation is an image processing operation based on mathematical morphology. It uses structuring elements such as rectangles and circles to expand the foreground area of an image, connecting adjacent foreground pixels and filling small holes. Its purpose is to strengthen the continuity of the foreground area, repair broken boundaries caused by noise, and provide complete regional information for subsequent contour detection.
[0094] Image binarization and morphological dilation operations achieve accurate separation and boundary enhancement of diagnostic areas through adaptive threshold segmentation and morphological restoration.
[0095] S114 locates the thyroid diagnosis area in the image by contour detection, and crops to obtain an image of the thyroid diagnosis area.
[0096] Contour detection refers to finding continuous, closed boundaries in a binary image and extracting the geometric contours of the foreground area. In this embodiment, it is used to accurately locate the actual boundary of the thyroid gland and exclude irrelevant areas in the image.
[0097] Contour detection can generate a mask image, which is a binary image with the same size as the original image, in which the thyroid area is marked as 255 (white) and the background is marked as 0 (black), which is used to accurately indicate the target area for subsequent cropping.
[0098] The thyroid region is cropped from the original image based on the mask image, and background noise such as the black edge of the device and surrounding tissue are excluded to generate an image containing only diagnostic information, providing high-quality input for the deep learning model.
[0099] S115 scales the image to a preset size and performs normalization processing.
[0100] The purpose of this step is to uniformly adjust the cropped thyroid B-ultrasound images to a fixed size, such as 256 × 256 pixels, to resolve the problem of size differences between different original images and meet the strict requirements of the subsequent deep learning model on the input dimension.
[0101] Differences in imaging brightness between different ultrasound devices can cause the model to mistakenly identify brightness as a pathological feature. Normalization is performed to unify the brightness baseline. Normalization scales the image pixel values from their original range (e.g., 0 to 255) to a near-standard normal distribution by subtracting the mean and dividing by the standard deviation.
[0102] In addition, in this embodiment, the images are uniformly saved in PNG format, which uses lossless compression to avoid the damage to medical details caused by lossy formats such as JPEG.
[0103] S12 uses transfer learning to train the EfficientNet-B0 model;
[0104] The EfficientNet-B0 model was trained using transfer learning. EfficientNet-B0 is a lightweight convolutional neural network that uses a compound scaling strategy to achieve a good balance between depth, width, and resolution. With only approximately 5M parameters, it outperforms the traditional ResNet-50 network. Its backbone architecture, consisting of 16 depthwise separable convolutional modules, effectively extracts multi-level features from images, from edge texture to semantic concepts. It is suitable for representing multi-scale information in thyroid B-ultrasound images, such as the local details of microcalcifications and the global structure of nodule morphology.
[0105] Before model training, the dataset used for training is first subjected to image enhancement processing, including random flipping, brightness perturbation, Gaussian blur and other operations, to improve the model's robustness to image diversity and enhance its generalization ability.
[0106] During training, the benign / malignant labels corresponding to B-ultrasound images were used as supervisory signals to train the parameters of the EfficientNet-B0 model. The training configuration was as follows: Adam was used as the optimizer, the initial learning rate was set to 1e-4, the batch size was 32, the loss function was Binary Cross Entropy Loss, and the total number of training epochs was 100. Early Stopping and the CosineAnnealingLR learning rate adjustment strategy were used to prevent overfitting.
[0107] S13 inputs the preprocessed image into the trained EfficientNet-B0 model to extract image features, thereby obtaining the image feature vector F US .
[0108] Specifically, to adapt to the multimodal feature analysis requirements of the present invention, the EfficientNet-B0 model is structurally adjusted, its classification head is removed, and only the backbone network is retained. The classification head is only used for specific classification tasks in the pre-training stage, and its output space is inconsistent with the task of distinguishing benign and malignant thyroid glands. If it is directly migrated without processing, it will affect the generalization ability of the features. After removing the classification head, the backbone network can be used as a general feature extractor to output high-dimensional, abstract image semantic features, which is more suitable for subsequent feature fusion, analysis or customized classifier training.
[0109] In this embodiment, the EfficientNet-B0 model does not directly output the classification results, but outputs the image feature representation. Specifically, after the B-ultrasound image with the input shape of [3,224,224] is input into the backbone network, the output feature map with the shape of [320,7,7] is output, that is, the image feature vector F US , which can fully characterize the structural and semantic content of the image. This feature vector serves as one of the important inputs of the multimodal diagnosis module of the present invention, providing reliable basic image information for the subsequent Transformer fusion model.
[0110] Through the trained EfficientNet-B0, the B-ultrasound image with an input shape of [3,224,224] is converted into an image feature vector F of [320,7,7]. US ; used to represent its structural and semantic features, and this feature vector serves as an important input of the multimodal diagnosis module of the present invention.
[0111] S2 extracts structured features from the thyroid cytology images of the patient to be diagnosed and obtains the structured feature vector FCYTO.
[0112] This step aims to use computer vision technology to extract structured cytological features from whole-slide imaging (WSI) images generated from thyroid cell samples obtained by fine needle aspiration (FNA). It specifically includes the following sub-steps:
[0113] S21 uses the YOLOv5 target detection model to train and detect cytology WSI images. The process is as follows:
[0114] S211 uses a sliding window strategy to crop the WSI image and generate image tiles of fixed size.
[0115] Since WSI images have extremely high resolution and huge overall size, and the cell targets are small and densely distributed, in order to achieve efficient processing and fine-grained analysis, the present invention adopts a sliding window strategy to crop the WSI images into several image blocks of fixed size, ensuring that the detection network can focus on the cell targets in the local area, thereby improving detection accuracy and speed.
[0116] S212 combines the labeled data provided by pathology experts to construct a cell detection dataset for training.
[0117] Based on the morphological characteristics of different types of cells in thyroid cytology images, pathologists with clinical experience were invited to annotate the cropped images category by category, including: mitotic cells, atypical cell nuclei, papillary carcinoma characteristic cells, follicular carcinoma characteristic cells, and other background or non-specific cells (such as macrophages, impurities, etc.), to construct a high-quality detection dataset as a supervisory signal for subsequent model training.
[0118] S213 uses the YOLOv5 model as the target detection network structure, and performs model training and medical adaptation.
[0119] This paper uses the YOLOv5 object detection framework as a cell target detection model. YOLOv5 boasts lightweight models, fast inference speed, and excellent small target recognition performance, making it suitable for detecting small-sized targets in cytological images. During model training, several optimization settings are implemented based on the characteristics of medical images, including using a smaller anchor matching threshold, reducing the learning rate, and balancing the weights of different target categories, to improve the model's detection and classification performance for multiple cell target types.
[0120] After training, the YOLOv5 model can identify and count multiple cell types in an image, including but not limited to: mitotic cells, atypical nuclei, papillary carcinoma-like cells, follicular carcinoma-like cells, and other atypical or background cells (such as macrophages and impurities).
[0121] The detection results output by the model can be quantified and statistically analyzed to form a structured feature vector F CYTO It represents the distribution characteristics of multiple types of cells in the patient's cytological images, providing a quantitative basis at the cell level for subsequent multimodal diagnostic models.
[0122] S22 uses the trained YOLOv5 target detection model to detect and count the multiple types of cells in the thyroid cytology WSI image, extract the structural information and generate the structural feature vector F CYTO .
[0123] After the detection is completed, further statistical analysis is performed on the detection results of each image to extract the structural statistical features of the cytological image, including but not limited to:
[0124] (1) Cell type statistics and order of magnitude characteristics are used to measure the abundance of a certain type of key cells in the sample and reflect the distribution of morphological characteristics of the lesions, including the number of mitotic cells, the number of atypical nuclei (nuclear atypia), the number of papillary carcinoma characteristic cells, the number of follicular carcinoma characteristic cells, and the number of other atypical or background cells such as macrophages and impurities.
[0125] (2) The average detection score is used to reflect the average judgment strength of the model for this type of target, and the confidence level of the model in identifying this type of cell, which can indirectly measure the clarity of the lesion.
[0126] The mean score of the detection box of each type of cell indicates the average confidence level of the target in the entire WSI.
[0127] (3) Average confidence probability statistics: The average value after the model output is softmaxed, which represents the consistency of the model judgment of the cell category. The closer the average confidence probability of each target prediction category is to 1, the more concentrated the prediction is.
[0128] (4) The median test score and median reliability statistics are less affected by extreme values and are suitable for describing the stability and representativeness of cell detection results.
[0129] All the above statistical features are aggregated per patient, and finally a structured feature vector F with a dimension of 45 is formed through a one-hot encoder. CYTO .
[0130] It should be noted that in this embodiment, the execution order of steps S1 and S2 is not limited and they can be performed simultaneously.
[0131] S3 takes the B-ultrasound image feature vector F US and cytological structured feature vector F CYTO Convert to a fusion token sequence.
[0132] This step aims to convert feature vectors from different modalities into a unified format and construct a fused token sequence to provide basic input for subsequent multimodal feature interaction and joint modeling. It specifically includes the following two sub-steps:
[0133] S31 converts the B-ultrasound image feature vector F US Convert to image token sequence token US
[0134] First, F UThe image feature map represented by
[15] is divided into n spatial regions through adaptive average pooling to obtain n image feature vector units. Each feature unit is then linearly transformed and mapped to a p-dimensional vector space to generate an image token, and modal category code and position code are attached to each token.
[0135] Specifically, in this embodiment:
[0136] The feature map is divided into n = 9 spatial regions, and 9 image feature vector units are extracted; each feature vector is projected into a p = 512-dimensional vector space through linear transformation to form 9 image tokens; linear transformation is implemented through matrix multiplication to unify the feature dimensions of each modality and solve the information fusion barrier caused by inconsistent original feature dimensions; modality category encoding adopts the form of learnable vectors, and the identification token comes from the ultrasound image, which helps the model distinguish the semantic differences between different modalities; position encoding represents the relative position of the image feature unit in two-dimensional space, assisting the model in capturing spatial structure and distribution patterns.
[0137] S32: Cytological structured feature vector F CYTO Convert to structured Token sequence Token CYTO
[0138] F CYTO The cells are divided into m structured feature units according to their cell types. Each unit contains multiple statistical features (such as quantity, average score, confidence, etc.). Each unit is then linearly transformed and mapped to a p-dimensional space to form a structured Token sequence. Similarly, modal category coding and position coding are added to each Token.
[0139] In specific implementation:
[0140] Each cell type (e.g., mitotic, nuclear atypia, PTCA, etc.) corresponds to a 5-dimensional statistical feature vector (number, mean score, mean confidence, median score, median position confidence); a total of $m=9$ cell types are divided, resulting in 9 structured feature vector units; each unit is linearly mapped to a 512-dimensional vector space, generating 9 structured tokens; the modality category code indicates that the token originates from the cytological image to avoid semantic confusion with the image token;
[0141] Positional encoding identifies the order in which cell types are arranged in the sequence. Because different cell types have potential priorities in clinical practice (e.g., suspected cancer cells are more important than background cells), their relative positions in the token sequence have certain biological implications.
[0142] S33 stitching image token USand structured tokens CYTO , build a fusion Token sequence Token fusion .
[0143] The n image tokens obtained in step S31 and the m structured tokens obtained in step S32 are concatenated in sequence to form a fused token sequence with a length of n+m and a dimension of p for each token, which serves as the input of the subsequent multimodal Transformer network.
[0144] Specifically:
[0145] Image Token US Sequence representation is derived from local area features of ultrasound images; structured token CYTO Sequence representation is derived from the statistical features of cytological images; the splicing order is usually image token US First, structured token CYTO After that, or according to the task requirements, you can customize the sorting; the fusion token obtained after splicing fusion The sequence is:
[0146]
[0147] in, Represents the i-th image Token, represents the jth structured token, with all tokens having 512 dimensions. This sequence will be input into the subsequent Transformer structure for cross-modal information fusion and global modeling. This operation completes the unified encoding of multimodal features and provides a standardized sequence input for subsequent diagnostic decision models.
[0148] S4: Classification prediction based on a multimodal prediction network of the Transformer architecture, outputting the probability of thyroid tumor malignancy
[0149] In this step, the fusion token fusion The sequence is input into a multimodal prediction network built on the Transformer architecture to complete the probability prediction of the patient's thyroid tumor being malignant. The specific steps are as follows:
[0150] S41: Multimodal Fusion Coding
[0151] The ultrasound image token US and Cytology Image Structured Token CYTO The spliced fusion sequence token fusion , which is input into the multimodal feature encoder composed of a multi-layer Transformer encoder stack.
[0152] The Transformer encoder consists of the following submodules:
[0153] Multi-head self-attention mechanism: Each layer contains 8 attention heads, which divide the input token sequence into multiple subspaces. These subspaces are independently modeled and then concatenated in different attention heads to learn the multiple dependencies between features of different modalities.
[0154] Some attention heads can capture the spatial structure between ultrasound image patches, such as the texture difference between the center and edge of a nodule; other attention heads can mine the logical relationship between cytological units, such as the ratio characteristics of lymphocytes and cancer cells;
[0155] Feedforward neural network (FeedForward) sublayer: This layer consists of two fully connected networks with a ReLU activation in the middle and a width of 1024. This layer enhances nonlinear expression capabilities and improves the modeling capabilities of key features (such as the combination of ultrasound edge irregularities and nuclear atypia).
[0156] Residual connections and LayerNorm normalization mechanism: Residual connections alleviate the vanishing gradient problem in deep network training, ensuring that low-level image texture features and structured statistical features can be effectively transmitted; LayerNorm improves the model's robustness to differences in the distribution of features across different modalities;
[0157] In this example, two layers of Transformer encoders are stacked:
[0158] The first layer mainly models the interaction of local features within the modality, such as the local texture consistency of image patches or the proportional relationship between cell groups;
[0159] The second layer further captures cross-modal global interaction features, such as the diagnostic correlation between the overall features of ultrasound images and cell morphology features.
[0160] S42: Fusion feature generation
[0161] The Token output by the Transformer encoder ' fusion The sequence (length n+m, dimension p) is averaged and averaged over the sequence dimension to obtain a fused feature vector of fixed length p.
[0162] Specifically:
[0163] Token ' fusion When the sequence length is 18 and the dimension is 512, average pooling is performed to generate a fused feature vector with a dimension of 512. This vector aggregates the deep interactive features of ultrasound patches and cytology units and has global diagnostic semantic information.
[0164] S43: Classification prediction
[0165] Input the fused feature vector into a classifier containing a two - layer fully - connected network in sequence to complete classification prediction:
[0166] The first - layer fully - connected network (hidden layer): Perform a linear transformation on the fused feature vector, compressing the dimension from 512 to 256; Through dimension compression, feature compression and redundancy filtering are achieved, highlighting features highly relevant to thyroid malignancy diagnosis (such as microcalcifications, mitotic figures, etc.), and preventing model over - fitting;
[0167] The second - layer fully - connected network (output layer): Further linearly map the 256 - dimensional feature to 1 - dimensional; Use the Sigmoid activation function to convert the output into a probability value in the range of [0, 1],
[0168] representing the predicted probability that the patient has thyroid malignancy.
[0169] Traditional binary classification models such as support vector machines output hard labels of either 0 or 1, which cannot reflect the confidence of the diagnosis. The probability value output by Sigmoid provides a more detailed risk assessment, for example:
[0170] 0.5 < P < 0.7: It is suggested that further examinations (such as puncture biopsy) are needed;
[0171] P < 0.3: It is possible to safely follow - up and observe, reducing unnecessary invasive operations.
[0172] In this embodiment, the classifier uses the AdamW optimizer with an initial learning rate of 1e - 4, combined with a cosine annealing learning rate scheduler; The Batch size is 16, the total number of training epochs is set to 100, and an Early Stopping mechanism is set to avoid over - fitting; The model training adopts a five - fold cross - validation method, and the evaluation metrics include accuracy, AUC, and F1 - score.
[0173] Among them, AUC represents the area under the Receiver Operating Characteristic curve, which is used to measure the comprehensive performance of the model's classification effect at different thresholds. The value range of AUC is usually between 0 and 1, where 1 represents a perfect classifier and 0.5 represents a random classifier. A higher AUC value indicates that the model can better distinguish positive and negative examples at different thresholds.
[0174] Accuracy is an index to measure the overall prediction correctness of the model, which considers the number of true positive (TP) and true negative (TN). The higher the accuracy, the better the overall performance of the model.
[0175] Accuracy=(TN+TP) / (TN+TP+FN+FP);
[0176] Precision measures the accuracy of the model in predicting positive examples, that is, the proportion of true positives (TP). High precision means that the model can accurately identify positive examples and reduce the number of false positives (FP).
[0177] Precision = TP / (TP+FP);
[0178] Specificity measures the accuracy of a model in predicting negative examples, that is, the proportion of true negatives (TN). A high specificity indicates that the model can accurately exclude negative examples and reduce the number of false positives (FP).
[0179] Specificity = TN / (FP + TN);
[0180] Sensitivity / Recall measures the model's ability to correctly identify positive examples, that is, the proportion of true positives (TP). High sensitivity means the model can capture more positive examples and reduce the number of false negatives (FN).
[0181] TP, FP, TN, and FN are true positive, false positive, true negative, and false negative, respectively.
[0182] The F1 score is the harmonic mean of precision and recall, which comprehensively considers the accuracy and comprehensiveness of the model. The F1 score is often used to balance the trade-off between precision and recall in different tasks.
[0183] This method employs fine-grained segmentation and unified encoding of features from different modalities, then performs deep fusion modeling within a Transformer architecture. This effectively captures the complex dependencies and diagnostic complementarity between ultrasound images and cytological features. Compared to single-modality models or simple splicing and fusion approaches, this architecture demonstrates higher classification accuracy and generalization in the task of classifying benign and malignant thyroid tumors.
[0184] Experimental results show that the Transformer structure used in this technical solution can fully model the contextual dependencies between different modal features and achieve deep semantic fusion. It is significantly superior to traditional feature splicing and fusion methods and single-modal models in the task of classifying benign and malignant thyroid tumors.
[0185] like Figure 2As shown in Figure 2, the proposed method was validated on actual clinical samples. Compared with single-modality prediction methods, the fusion model achieved significant improvements in diagnostic accuracy, AUC, sensitivity, and specificity. Furthermore, comparative experiments have shown that the Transformer fusion architecture outperforms traditional feature concatenation methods and can better model the interactions between heterogeneous modalities.
[0186] Example 2:
[0187] This embodiment provides a thyroid tumor diagnosis system for combined analysis of ultrasound and cytological images, which includes:
[0188] The B-ultrasound image feature vector extraction module is used to convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US ;
[0189] The cytology image structured feature extraction module is used to extract the structured features of the thyroid cytology image of the patient to be diagnosed and obtain the structured feature vector F CYTO ;
[0190] Feature vector fusion module, used to fusion the image feature vector F US And the structured feature vector F CYTO Convert to fusion token sequence;
[0191] The prediction module is used to input the fused token sequence into the multimodal prediction network based on the Transformer architecture for classification prediction, thereby obtaining the probability that the thyroid tumor of the patient to be diagnosed is malignant.
[0192] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0193] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a computer terminal (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0194] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.
Claims
1. A method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images, characterized in that include: Convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US ; The structural feature F is extracted from the thyroid cytology image of the patient to be diagnosed. CYTO ; The B-ultrasound image feature vector F US and cytological structured feature vector F CYTO Convert to fusion token sequence; The fused token sequence is encoded and input into the classification model for classification prediction, thereby obtaining the probability that the thyroid tumor of the patient to be diagnosed is malignant.
2. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 1, characterized in that Convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US The steps include: Preprocessing of thyroid B-ultrasound images of patients to be diagnosed; The preprocessed image is input into the trained EfficientNet-B0 model for image feature extraction to obtain the image feature vector F US .
3. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 2, characterized in that The steps for preprocessing the thyroid ultrasound images of the patient to be diagnosed include: Crop the image edges according to the preset ratio; Perform white balance processing and grayscale conversion; The image is converted into a single-channel grayscale image through image binarization and morphological dilation operations; The thyroid diagnostic area is located in the image by contour detection, and the thyroid diagnostic area image is obtained by cropping; Scale the image to a preset size and normalize it.
4. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 2, characterized in that: Before training the EfficientNet-B0 model, image augmentation is performed on the dataset used for training; During training, the benign / malignant labels are used as supervisory signals to train the EfficientNet-B0 model; The training parameters include: Adam optimizer, initial learning rate 1e-4, batch size 32, binary cross entropy loss function, 100 training rounds, early stopping mechanism and CosineAnnealingLR learning rate adjustment strategy. In the feature extraction stage, the classification head in the EfficientNet-B0 model is removed, and only the backbone network is retained.
5. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 1, characterized in that: Use cytology WSI images to train the YOLOv5 object detection model; The trained YOLOv5 target detection model is used to detect and count the multiple types of cells in the thyroid cytology WSI images of the patients to be diagnosed, thereby extracting structural information and generating a structural feature vector F. CYTO ,include: Use the trained YOLOv5 object detection model to detect and classify cells in thyroid cytology WSI images, perform statistical analysis based on the detection results, and extract structured statistical features of the images; The structured statistical features include at least one of cell type statistics, average detection score statistics, average confidence statistics, median detection statistics, and median position confidence statistics; Aggregate the above statistical features based on patients to generate a fixed-dimensional structured feature vector F CYTO .
6. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 5, characterized in that: The steps for training the YOLOv5 object detection model include: A sliding window strategy is used to crop the WSI image to generate image patches of fixed size; Using the cell category and location annotation data provided by pathology experts, a cell detection training dataset was constructed to train the YOLOv5 object detection model; The YOLOv5 target detection model can detect and identify mitotic cells, atypical cell nuclei, papillary carcinoma characteristic cells, and follicular carcinoma characteristic cells.
7. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 1, characterized in that The B-ultrasound image feature vector F US and cytological structured feature vector F CYTO The steps for converting to a fusion token sequence include: The image feature vector F US The feature map represented is divided into n spatial regions through adaptive average pooling, obtaining n image feature vector units; each unit is mapped to a p-dimensional vector space through linear transformation to form n image tokens; the corresponding modality category code and position code are added to each image token; The structured feature vector F CYTO Divide into m structured feature vector units according to cell type, and map each unit to p-dimensional vector space through linear transformation to form m structured tokens; modal category code and position code are also added to each structured token; The n image tokens are concatenated with the m structured tokens to form a fused token sequence with a length of n+m and a dimension of p.
8. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 1, characterized in that: The steps of classification prediction include: The fused token sequence is fed into a multimodal encoder consisting of a stack of multiple layers of Transformer encoders, which include a multi-head self-attention mechanism, a feedforward neural network sublayer, residual connections, and a normalization mechanism. The Token sequence output by the multimodal encoder is subjected to an average pooling operation to generate a fused feature vector; The fused feature vector is input into a classifier comprising a two-layer fully connected network to sequentially complete feature compression and malignancy probability prediction to obtain the predicted probability that the thyroid tumor of the patient to be diagnosed is malignant.
9. The method for diagnosing thyroid tumors by combined analysis of ultrasound and cytological images according to claim 8, characterized in that: The classifier uses the AdamW optimizer with an initial learning rate of 1e-4, combined with a cosine annealing learning rate scheduler; a batch size of 16 is used, the total number of training rounds is set to 100 rounds, and an early stopping mechanism is set to avoid overfitting; the model training adopts a five-fold cross-validation method, and the evaluation indicators include accuracy, AUC, and F1 score.
10. A thyroid tumor diagnosis system that combines ultrasound and cytological image analysis, characterized in that include: The B-ultrasound image feature vector extraction module is used to convert the thyroid B-ultrasound image of the patient to be diagnosed into the image feature vector F US ; The cytology image structured feature extraction module is used to extract the structured features of the thyroid cytology image of the patient to be diagnosed and obtain the structured feature vector F CYTO ; Feature vector fusion module, used to fusion the image feature vector F US And the structured feature vector F CYTO Convert to fusion token sequence; The prediction module is used to encode the fused token sequence and input it into the classification model for classification prediction, thereby obtaining the probability that the thyroid tumor of the patient to be diagnosed is malignant.
Citation Information
Cited By
AI-based traditional Chinese medicine nursing recommendation system
CN120878096A
Malignant nodule detection image processing method based on benign thyroid
CN121190423A
Domain-adaptive thyroid ultrasound image typing method and device
CN121921302A
Thyroid nodule BRAF gene mutation prediction system and method based on multi-modal fusion
CN122157768A
ANA cell image multi-mode analysis and prediction method, system, equipment and medium
CN122176705A