A classification method for benign and malignant thyroid nodules based on segmentation-based multi-feature information

By using a multi-scale segmentation network and a knowledge-guided multi-branch classification network on the ultrasound image of thyroid nodules, the problem of neglecting knowledge in the field of medical diagnosis in the prior art is solved, and the accuracy of benign and malignant classification of thyroid nodules is significantly improved.

CN117911772BActive Publication Date: 2025-05-06INNER MONGOLIA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410083374.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-19
Publication Date
2025-05-06
Estimated Expiration
2044-01-19

AI Technical Summary

Technical Problem

Existing AI diagnostic models process medical images the same as other images, ignoring the key domain knowledge of medical diagnostic tasks, resulting in poor classification of benign and malignant thyroid nodules.

Method used

The benign and malignant classification method of thyroid nodules based on segmentation multi-feature information is adopted, and the ultrasound images of thyroid nodules are segmented and classified through multi-scale segmentation networks and knowledge-guided multi-branch classification networks are segmented and classified. The method includes data set processing, classification network construction, model training, and nodule benign and malignant classification.

Benefits of technology

The multi-scale characteristics of thyroid nodules are obtained through segmentation networks, combined with the knowledge guidance mechanism of the multi-branch classification network, the accuracy of benign and malignant classification of thyroid nodules is significantly improved, and the problem of neglecting domain knowledge in the existing technology is overcome.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117911772B_ABST
    Figure CN117911772B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for classifying benign and malignant thyroid nodules based on segmentation-based multi-feature information, which belongs to the field of medicine and information technology, and includes the following steps: step S1: data set processing; step S2: classification network construction; step S3: classification model training; step S4: benign and malignant nodule classification. According to expert experience, the present invention classifies thyroid nodule ultrasound images into benign and malignant types by adopting a segmentation + classification strategy; in the segmentation network, the AG attention mechanism is used to obtain the spatially accurate information of low-level features, and the information of irrelevant areas is suppressed to reduce redundancy; in order to overcome the grid problem and sampling sparseness problem of dilated convolution, the DASPP module is designed; finally, the network's expression ability is enhanced by multi-scale information fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medicine and information technology, and in particular to a method for classifying benign and malignant thyroid nodules based on segmentation of multi-feature information. Background Art

[0002] Ultrasonic examination uses the physical properties of ultrasound and the acoustic parameters of human tissue to perform imaging, combining medical imaging techniques from disciplines such as anatomy, pathophysiology and clinical medicine. Sound waves with a frequency exceeding the upper limit of human hearing, 20kHZ, are called ultrasound. Medical ultrasound has multiple physical properties such as directionality, reflection, refraction, scattering, attenuation, absorption and Doppler effect. When sound waves pass through the interfaces of different tissues and organs, they form echoes of different intensities. These echo signals are processed by a computer and then imaged.

[0003] Ultrasound examination is currently a routine screening method for thyroid nodules, and computer-aided diagnosis can provide doctors with objective advice. However, existing artificial intelligence diagnostic models treat medical images in the same way as other images, ignoring the key domain knowledge of medical diagnostic tasks. To this end, a method for classifying benign and malignant thyroid nodules based on segmentation-based multi-feature information is proposed. Summary of the invention

[0004] The technical problem to be solved by the present invention is: how to solve the problem that the existing artificial intelligence diagnosis model processes medical images in the same way as other images, ignoring the key field knowledge of medical diagnosis tasks, and provides a method for classifying benign and malignant thyroid nodules based on segmentation of multi-feature information.

[0005] The present invention solves the above technical problems through the following technical solutions, and the present invention comprises the following steps:

[0006] Step S1: Dataset processing

[0007] Obtain multiple thyroid nodule images, annotate benign and malignant labels under the guidance of ultrasound doctors to form a data set, perform data enhancement processing on the images in the data set, and then divide the data set into training set, test set and validation set;

[0008] Step S2: Classification network construction

[0009] Constructing a benign or malignant thyroid nodule classification network, which includes a multi-scale segmentation network and a knowledge-guided multi-branch classification network. The multi-scale segmentation network is used to segment the thyroid nodules in the image to obtain a nodule mask image, and the multi-branch network is used to use the segmented nodule region image, nodule edge image, and nodule original image as input for classification.

[0010] Step S3: Classification model training

[0011] The training set is used to train the benign and malignant classification network of thyroid nodules to obtain a trained benign and malignant classification model of thyroid nodules;

[0012] Step S4: Classification of benign and malignant nodules

[0013] The benign and malignant thyroid nodule classification model is verified in the test set, and the benign and malignant thyroid nodule classification model that has passed the verification is used to classify the benign and malignant thyroid nodules in the image to be detected to obtain the benign and malignant nodule classification results.

[0014] Furthermore, in the step S1, when marking the benign or malignant labels, the nodule contour is outlined to generate a mask image; the data enhancement processing is to rotate the images in the data set horizontally and vertically to expand the data set.

[0015] Furthermore, in the step S2, the multi-scale segmentation network uses the U-Net network as the main network, including a DASPP module, an attention gate and a multi-scale feature fusion module; wherein the attention gate is used to obtain the spatially accurate information of the low-level features and suppress the information of irrelevant areas; the DASPP module is used to splice the branches with low expansion coefficients with the branches with high expansion coefficients in a cascade manner and then perform a convolution operation; the multi-scale feature fusion module is used to fuse the output information of the decoder at different scales to obtain a multi-scale feature representation, and then upsample the features of all scales to restore the original size, and use a 1×1 convolution to reduce the channel dimension to 1, and finally splice them for output, so as to obtain a segmentation result.

[0016] Furthermore, in step S2, the multi-branch classification network includes three branch networks, and the three branch networks have the same structure, namely, an original branch network, a regional branch network, an edge branch network and a cross-level feature fusion module; wherein, the original branch network is used to extract the location information and size information of the nodule, and its input is the original image of the nodule, the regional branch network is used to extract the internal information of the nodule and the aspect ratio information of the nodule, and its input is the nodule regional image, the edge branch network is used to extract the edge information of the nodule, and its input is the nodule edge image, and the cross-level feature fusion module is used to fuse the information extracted by the original branch network, the regional branch network, and the edge branch network, and then classify the benign and malignant thyroid nodules.

[0017] Furthermore, in step S2, the process of acquiring the nodule region image and the nodule edge image is as follows:

[0018] S21: cropping the original nodule image using the nodule mask image obtained by segmentation to obtain a square nodule region image;

[0019] S22: then determining the edge of the nodule, and expanding the inner and outer edges at equal distances according to the edge to obtain a mask image of the nodule edge;

[0020] S23: Crop the nodule original image according to the nodule edge mask image to obtain a nodule edge image.

[0021] Furthermore, the three branch networks are all improved ResNet50 feature extraction networks, and the CA attention mechanism is used in each residual module of the improved ResNet50 feature extraction network. The CA attention mechanism is used to pay attention to the input features in the horizontal and vertical directions and act on the input. Each element in the two-directional attention indicates whether there is an area of ​​interest in the corresponding row and column. The output features of different residual modules are spliced ​​and fused in each branch of the improved ResNet50 feature extraction network.

[0022] Compared with the prior art, the present invention has the following advantages: the method for classifying benign and malignant thyroid nodules based on segmentation-based multi-feature information classifies benign and malignant thyroid nodules in ultrasound images according to expert experience by adopting a segmentation + classification strategy; in the segmentation network, the AG attention mechanism is used to obtain accurate spatial information of low-level features, and suppress information in irrelevant areas to reduce redundancy; in order to overcome the grid problem and sampling sparseness problem of dilated convolution, a DASPP module is designed; finally, the expression ability of the network is enhanced through multi-scale information fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a schematic diagram of the structure of a benign and malignant thyroid nodule classification model according to an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of the structure of a multi-scale segmentation network (i.e., a multi-scale attention gate U-Net network) in an embodiment of the present invention;

[0025] Figure 3 is a schematic diagram of the structure of a multi-branch classification network in an embodiment of the present invention;

[0026] FIG4( a ) is an example of a benign nodule image in an embodiment of the present invention;

[0027] FIG4( b ) is an example of a malignant nodule image in an embodiment of the present invention;

[0028] Figure 5 Schematic diagram of the process of obtaining a nodule edge image in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following is a detailed description of an embodiment of the present invention. This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given, but the protection scope of the present invention is not limited to the following embodiment.

[0030] This embodiment provides a technical solution: a method for classifying benign and malignant thyroid nodules based on segmented multi-feature information, including the following main steps:

[0031] Step 1: Data acquisition and preprocessing

[0032] The data set used in this embodiment comes from the ultrasound department of a certain regional people's hospital, with a total of 4021 thyroid nodule images (ultrasound images), including 1844 benign nodule images and 2177 malignant nodule images. All ultrasound images have been through ethical review and data desensitization, and under the guidance of experienced ultrasound doctors, benign and malignant labels have been marked, and the outline of the nodule has been completed to generate a mask image. These data have been desensitized and only include thyroid nodule images. The nodule ultrasound image data is shown in Figure 4 (a) and Figure 4 (b). At the same time, the data set is divided into a training set, a test set, and a validation set without an intersection according to a ratio of 3: 1: 1. And the images in the data set are rotated horizontally and rotated up and down to expand the data set.

[0033] Step 2: Establish a classification model for benign and malignant thyroid nodules guided by clinical knowledge

[0034] like Figure 1 As shown in the figure, the thyroid nodule benign and malignant classification model is mainly divided into two parts, namely the nodule segmentation network and the nodule classification network. First, a multi-scale segmentation network is designed to obtain the nodule mask image. Then, a multi-branch classification network is constructed for nodule classification based on the clinical diagnosis experience of doctors.

[0035] 1. Multi-scale segmentation network

[0036] Thyroid nodules vary in size, and small thyroid nodules account for a very high proportion in ultrasound images. However, there are many difficulties in segmenting small objects. Compared with objects of other sizes, small objects have fewer pixels, making it difficult to extract effective features, and are easily disturbed by the background, resulting in inaccurate segmentation.

[0037] As the segmentation network deepens, the more contextual information can be learned, and the stronger the feature expression ability. However, it will also cause the high-order features to lose a lot of spatial precision information, and at the same time there is a lot of redundant information in the low-level features. This makes it difficult to improve the segmentation accuracy of small thyroid nodules. To solve this problem, the attention gate (AG) is used to obtain the relevant spatial precision information of the low-level features, and suppress the information of irrelevant areas to reduce redundancy. In addition, the dilated convolution also has a gridding issue. When the dilation coefficient is very large, the sampling of the input will become sparse, resulting in the loss of some information and some information over long distances may not be relevant. In addition, the gridding problem may interrupt the continuity between local information. To overcome this problem, the improved ASPP module, Densely Connected Atrous Spatial Pyramid Pooling (DASPP), is integrated into the network. It uses a cascade method to splice the branch with a low expansion coefficient with the branch with a high expansion coefficient and then perform a convolution operation. This not only obtains the scale information of the original image, but also can further extract information based on the information extracted at other scales. Multiple sampling not only increases the receptive field of the feature, but also alleviates the grid problem of dilated convolution, which is beneficial to the segmentation of small targets to a certain extent. Finally, after upsampling the semantic information of different scales to restore it to its original size, they are fused to obtain multi-scale semantic information, further improving the segmentation effect of small nodules. The segmentation network uses the U-Net network as the backbone network. Based on the above problems, a multi-scale attention gate U-Net (Multi-scale attention gate U-Net, MS-AGU-Net) network is established. Its network structure is as follows: Figure 2 As shown:

[0038] 1) Attention Gate (AG)

[0039] The attention mechanism is similar to human visual attention, focusing on certain things and ignoring other things. AG can suppress irrelevant information while enhancing useful information.

[0040] 2)Densely Connected Atrous Spatial Pyramid Pooling(DASPP)

[0041] Dilated convolution is the basis of the ASPP module. Compared with ordinary convolution, dilated convolution sets a dilation coefficient to increase the distance between pixels in each row and column of the convolution kernel, thereby increasing the receptive field of the convolution kernel.

[0042] ASPP obtains multi-scale information from multiple scales by connecting multiple dilated convolutions with different dilation coefficients in parallel. Multi-scale information allows the model to construct spatial position relationships under different scale perspectives for the target, which can help the model better locate key features and facilitate more comprehensive judgment and recognition.

[0043] 3) Multi-scale feature fusion module

[0044] The multi-scale feature fusion module is used to fuse the output information of different scales of the decoder to obtain multi-scale feature representation, so that the model can obtain information of different scales to enhance the model accuracy. The features of all scales are upsampled to restore the original size, and the channel dimension is reduced to 1 using 1×1 convolution, and finally they are spliced ​​and output to obtain the prediction result.

[0045] 2. Multi-branch classification network

[0046] In the TIRADS classification guidelines, important ultrasound features such as nodule composition, internal echo, aspect ratio, edge smoothness, and edge shape are the basis for distinguishing nodule grade and benign or malignant. Based on expert knowledge, a three-branch classification network was designed, using three complementary branches to learn multi-view features of different regions. The original image of the thyroid nodule, the nodule region image, and the nodule edge image were input into the three branch networks respectively to obtain global features, nodule internal features, and nodule edge features, respectively, to enhance the classification effect. The classification network structure diagram is shown above Figure 3 As shown, the structures of all branch networks are the same.

[0047] In the classification network, the original ultrasound image is cropped into nodule region images and nodule edge images according to the segmented image. They are input into the knowledge-guided three-branch classification network together with the original image. The purpose of this is to allow the original image to obtain global information, and it can also play a supplementary role when the segmentation effect is not ideal. At the same time, the nodule region image obtains the internal information of the nodule, and the nodule edge image obtains the nodule edge information. CA attention mechanism and cross-level feature fusion are used in the classification network. The CA attention mechanism is used to obtain channel dimension and spatial dimension information to enhance the feature extraction ability of the network. Finally, the information of different scales is fused through cross-level feature fusion to obtain the final classification feature, thereby realizing the classification of benign and malignant nodules.

[0048] 1) Network architecture

[0049] Usually CNN requires an input of a standard size, such as 224×224. This requires some processing of the original image to meet the input requirements. However, the images obtained are rarely square, and most of them are rectangular. The model will randomly crop or compress the input image, which will destroy the information contained in the original image. In order to maintain the true aspect ratio information of different nodules, according to the nodule area obtained by segmentation, each nodule of size H×W is cropped with the maximum value of its length and width as the side length to make it a square. For some images that cannot be cropped into squares, the smallest side is padded with zeros on the basis of the cropping to make it a square. In this way, the nodule will not be compressed in one direction, and its original aspect ratio information can be maintained.

[0050] The input of the first branch is the original image of the nodule, which allows the model to capture the location and size information of the nodule. In addition, the characteristics of the tissues surrounding the nodule will also affect the nodule classification, such as diffuse sclerosis and the difference in internal and external echoes. At the same time, when the segmentation effect is not ideal, this branch can also assist and supplement the classification results.

[0051] The input of the second branch is the cropped nodule region image. The nodule region is cropped into a square according to the segmented image, and then the image is only scaled without random cropping. In this way, the model can not only learn the information inside the nodule, but also the aspect ratio information of the nodule. In addition, the edge information of the nodule is also a very important feature in diagnosis, such as whether the edge is smooth and clear.

[0052] The input of the third branch is the nodule edge image. The edge image is cropped into a square using the same method and only scaled so that the network can extract more salient features.

[0053] In this way, it is necessary to crop the original nodule image according to the mask image obtained by segmentation, so as to obtain the nodule area image and the nodule edge image. The present invention processes the nodule image through OpenCV. OpenCV is an open source computer vision library that supports multiple programming languages ​​and can be used to process digital images and video images. It provides many functions, including functions ranging from filtering to object detection. The present invention mainly uses some of the target detection functions in OpenCV to detect the white area and the edge of the white area in the mask image. Figure 5 As shown, the original image of the nodule is firstly cropped by the mask image to obtain a square nodule region image. Then the edge of the nodule is determined, and the inside and outside are expanded at equal distances according to the edge to obtain a nodule edge mask image. Finally, the original image is cropped according to the edge mask image to obtain a nodule edge image.

[0054] The present invention uses the ResNet50 network as the baseline network of each branch network. In order to improve the classification performance of the network, two improvements are made to the network structure: (1) the output features of different residual modules are spliced ​​and fused in each branch of the classification network; (2) the Coordinate Attention (CA) attention mechanism module is used in each residual module of the classification network.

[0055] The attention mechanism module enhances the feature representation of the network by acquiring different types of information in the channel dimension or spatial dimension, thereby enhancing the feature extraction ability of the network. At present, most attention mechanisms globally pool the input features to obtain the global average or maximum value of the features. For example, the SE module encodes the features of each channel into a global feature through global average pooling, predicts the importance of each channel, or in other words, each channel obtains a different weight, and finally multiplies it with the original feature map, thereby enhancing or suppressing the channel information according to the importance of each channel. Although the SE module has achieved good results, the SE module only considers the channel information and ignores the spatial position information. In order to consider both channel information and spatial information at the same time, the CA attention mechanism is used in each Block (residual module) of the ResNet50 network.

[0056] The CA attention mechanism embeds the spatial dimension information into the channel dimension while considering the channel dimension information, so that the network can obtain information of a larger area. As mentioned above, the CA attention mechanism can perceive the input features in the horizontal and vertical directions and act on the input. Each element in the two-directional attention can indicate whether there is an area of ​​interest in the corresponding row and column. This method allows the network to better determine the specific location of the area of ​​interest, thereby improving the classification effect of the model.

[0057] In summary, the method for classifying benign and malignant thyroid nodules based on segmentation of multi-feature information in the above embodiment, according to expert experience, classifies benign and malignant thyroid nodules ultrasound images by adopting a segmentation + classification strategy; in the segmentation network, the AG attention mechanism is used to obtain the spatially accurate information of low-level features, and suppress the information of irrelevant areas to reduce redundancy; in order to overcome the grid problem and sampling sparseness problem of dilated convolution, the DASPP module is designed; finally, the expression ability of the network is enhanced through multi-scale information fusion.

[0058] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present invention.

Claims

1. A method for classifying benign and malignant thyroid nodules based on segmentation-based multi-feature information, characterized in that: The following steps are involved: Step S1: Dataset processing Obtain multiple thyroid nodule images, annotate benign and malignant labels under the guidance of ultrasound doctors to form a data set, perform data enhancement processing on the images in the data set, and then divide the data set into training set, test set and validation set; Step S2: Classification network construction Constructing a benign or malignant thyroid nodule classification network, which includes a multi-scale segmentation network and a knowledge-guided multi-branch classification network. The multi-scale segmentation network is used to segment the thyroid nodules in the image to obtain a nodule mask image, and the multi-branch network is used to use the segmented nodule region image, nodule edge image, and nodule original image as input for classification. In the step S2, the multi-scale segmentation network uses the U-Net network as the main network, including a DASPP module, an attention gate, and a multi-scale feature fusion module; wherein the attention gate is used to obtain accurate spatial information of low-level features and suppress information of irrelevant areas; the DASPP module is used to splice branches with low expansion coefficients with branches with high expansion coefficients in a cascade manner and then perform a convolution operation; the multi-scale feature fusion module is used to fuse output information of different scales of the decoder to obtain a multi-scale feature representation, and then upsample the features of all scales to restore the original size, and use a 1×1 convolution to reduce the channel dimension to 1, and finally splice them for output, so as to obtain a segmentation result; In step S2, the multi-branch classification network includes three branch networks, the three branch networks have the same structure, and the three branch networks are respectively an original branch network, a regional branch network, an edge branch network and a cross-level feature fusion module; wherein the original branch network is used to extract the location information and size information of the nodule, and its input is the original image of the nodule, the regional branch network is used to extract the information inside the nodule and the aspect ratio information of the nodule, and its input is the nodule regional image, the edge branch network is used to extract the edge information of the nodule, and its input is the nodule edge image, and the cross-level feature fusion module is used to fuse the information extracted by the original branch network, the regional branch network, and the edge branch network, and then classify the benign and malignant thyroid nodules; Step S3: Classification model training The training set is used to train the benign and malignant classification network of thyroid nodules to obtain a trained benign and malignant classification model of thyroid nodules; Step S4: Classification of benign and malignant nodules The benign and malignant classification model of thyroid nodules is verified in the test set, and the benign and malignant classification model of thyroid nodules is used to classify the benign and malignant nodules in the detected images to obtain the benign and malignant classification results of the nodules.

2. The method for classifying benign and malignant thyroid nodules based on segmented multi-feature information according to claim 1, characterized in that: In the step S1, when marking benign or malignant labels, the nodule contour is outlined to generate a mask image; the data enhancement processing is to perform horizontal rotation and vertical rotation processing on the images in the data set to expand the data set.

3. The method for classifying benign and malignant thyroid nodules based on segmented multi-feature information according to claim 1, characterized in that: In step S2, the process of acquiring the nodule region image and the nodule edge image is as follows: S21: cropping the original nodule image using the nodule mask image obtained by segmentation to obtain a square nodule region image; S22: then determining the edge of the nodule, and expanding the inner and outer edges at equal distances according to the edge to obtain a mask image of the nodule edge; S23: Crop the nodule original image according to the nodule edge mask image to obtain a nodule edge image.

4. The method for classifying benign and malignant thyroid nodules based on segmented multi-feature information according to claim 3, characterized in that: The three branch networks are all improved ResNet50 feature extraction networks. The CA attention mechanism is used in each residual module of the improved ResNet50 feature extraction network. The CA attention mechanism is used to perceive the input features in the horizontal and vertical directions and act on the input. Each element in the two-directional attention indicates whether there is an area of ​​interest in the corresponding row and column. The output features of different residual modules are spliced ​​and fused in each branch of the improved ResNet50 feature extraction network.