Medical image recognition and classification method based on multi-view integrated network
Through a multi-view integrated network, combining Curvelet transformation and depth separation convolution, the frequency and spatial domain characteristics of medical images are extracted, and the results of fusing multiple pre-trained models are combined through a multi-expert integration network, the subjectivity and limitations of feature design in traditional methods are solved, significantly improving the efficiency and accuracy of medical image classification.
Patent Information
- Application Number
- CN202510042243.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Traditional medical image recognition methods rely on artificially designed feature extraction and classification models, and have subjectivity and limitations of feature design, making it difficult to effectively improve diagnostic efficiency and accuracy.
A multi-view integrated network is adopted, and the frequency and spatial domain characteristics of the image are extracted through Curvelet transformation and depth separation convolution, and the results of multiple pre-trained models are fused through the multi-expert integration network to build a medical image classification model with strong robustness and excellent generalization ability.
It significantly improves the efficiency and accuracy of medical imaging classification, reduces dependence on feature design, improves the robustness and generalization capabilities of the model, and provides scientific and reliable auxiliary decision-making support for clinical practice.
Smart Images

Figure CN119992171A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of medical image detection, and in particular relates to a medical image recognition and classification method based on a multi-view integrated network. Background Art
[0002] With the rapid development of medical imaging technology, medical imaging detection has become an important means of modern medical diagnosis. As a common endocrine disease, the diagnosis of thyroid disease often relies on the auxiliary analysis of medical images such as ultrasound and CT. However, the traditional manual reading method is easily affected by the doctor's experience level and subjective judgment, which may lead to uncertainty in the diagnosis results. Therefore, the development of efficient and accurate automated classification methods is of great significance to improve diagnostic efficiency and accuracy.
[0003] Traditional machine learning methods in medical image recognition mainly rely on manually designed feature extraction and classification models. The typical process includes preprocessing the image (such as denoising, segmentation, etc.), extracting features based on texture, shape or statistics, and then completing the classification through classifiers such as support vector machine (SVM), random forest (RF) or k-nearest neighbor (k-NN). However, this method is highly dependent on features, and the quality of features directly affects the performance of the model. At the same time, feature design often requires the knowledge of domain experts, which is highly subjective and limited.
[0004] With the rapid development of deep learning technology, medical image analysis has entered a new stage of automation and intelligence. By combining a variety of advanced algorithms, deep learning methods can autonomously learn and extract key features, greatly improving recognition and classification efficiency and accuracy. Summary of the invention
[0005] The present invention fully exploits the image features of the region of interest and uses a multi-view integration strategy to fuse information from different perspectives and levels to build a classification model with strong robustness and excellent generalization ability. This technology can significantly improve the efficiency and accuracy of medical image classification, provide scientific and reliable auxiliary decision support for clinical practice, and has broad application prospects.
[0006] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:
[0007] A medical image recognition and classification method based on a multi-view integrated network comprises the following steps:
[0008] Collect and preprocess CT medical images to build a CT image dataset for training;
[0009] Construct a medical image classification model based on a multi-view integrated network, including: a C-DSC module based on Curvelet transform and deep separable convolution, and then connect it to a multi-expert integrated network; use CT image dataset data to conduct supervised training on the medical image recognition classification model, and select the model with the smallest classification error as the ideal classification model;
[0010] The unknown CT image is passed into the ideal classification model to obtain the recognition and classification results of the current image.
[0011] The pre-processing comprises the following steps:
[0012] Modify image data to a uniform size;
[0013] Manually mark the region of interest on the acquired CT image as the image label R;
[0014] Perform data augmentation on the annotated regions of interest to expand the data set and obtain multiple views C of N regions of interest of different sizes;
[0015] The original grayscale view O is superimposed, and finally N+1 multi-views are obtained.
[0016] The data enhancement method is:
[0017] The ROI area is cut out according to the marked box. Through the expansion operation, the center of the marked box is used as the reference, first expanding along the long side direction, then expanding along the wide side direction, each time expanding by 1 / 4 times, and each expansion is based on the previous step to obtain a new view.
[0018] In the C-DSC module, the data is subjected to Curvelet transformation to obtain two curvelet components (H, L) which are stacked with the original image to form a three-channel image. The spatial features of each channel are extracted through deep separable convolution, and then the grayscale, high-frequency and low-frequency features are effectively integrated.
[0019] The Curvelet transform is:
[0020] Convert it from the spatial domain to the frequency domain through Fast Fourier Transform (FFT);
[0021] The frequency domain data is rotated and resampled by setting random directions and angles to capture more diverse edge and directional features;
[0022] The frequency domain information is converted into the spatial domain through the inverse fast Fourier transform IFFT to obtain two curvelet components H and L; the high-frequency component H, the low-frequency component L and the original grayscale D image form a three-channel input X.
[0023] The depthwise separable convolution is:
[0024] Input X into 3*3 convolution to get the feature map of the same size as X ′ ; Then the size remains unchanged after the ReLU activation function;
[0025] After information fusion through 1*1 convolution, the feature map X1′ is obtained, which contains the weight information of each channel; after another ReLU activation function, the size of the feature map remains unchanged, and the feature map D is obtained.
[0026] The multi-expert integrated network includes a backbone network and a slave network, which is any one of DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet, and AlexNet networks pre-trained on ImageNet.
[0027] The multi-expert integrated network includes:
[0028] Input the multi-feature map D into the first N+1 slave networks with different freezing parameters, and input the image of the region of interest with the largest expansion scale into Expert Net6;
[0029] From the network, five Expert Net networks are used to train the views of N+1 (C and O) channels to be processed, and the output classification results are x1, x2, x3, x4, x5 respectively. The results are concatenated into feature map f1, and the Cat operation is performed to change the feature map from f1 to f2;
[0030] The result x6 output by the backbone network Expert Net6 is concatenated with f2 and subjected to Cat operation to become f3, which is output as the final recognition result.
[0031] The backbone network adopts the Efficientnet network with unfrozen parameters, and the slave network adopts the AlexNet, DenseNet, GoogleNet, ResNet, and ShuffleNet networks with frozen parameters; and the last fully connected layer of each network in the integrated network is set to 2 units.
[0032] A medical image recognition and classification system based on a multi-view integrated network comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement the medical image recognition and classification method based on a multi-view integrated network as described above when loading and executing the computer program.
[0033] The present invention has the following beneficial effects and advantages:
[0034] 1. The present invention is based on ensemble learning and uses C-DSC to enhance the feature expression ability of medical images. Through the combination of Curvelet transform and depthwise separable convolution, the data is converted between the frequency domain and the spatial domain. The model extracts key information from the frequency domain and the spatial domain respectively, performs feature fusion, enriches the feature extraction process, and improves the expression ability of the model.
[0035] 2. At the same time, the present invention combines AlexNet, DenseNet, GoogleNet, ResNet, ShuffleNet, and Efficientnet networks pre-trained on the ImageNet dataset to build an integrated network to train and classify the CT image dataset. It effectively solves the problem that medical images are difficult to capture complex multi-scale and multi-directional features, significantly improves the model's sensitivity to texture and edge information, and at the same time, due to the insufficient generalization ability of a single model, by integrating multiple pre-trained models and combining the division of labor strategy of the attending expert network and the auxiliary expert network, the model robustness and generalization ability are effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flow chart of the method of the present invention;
[0037] Figure 2 It is a schematic diagram of the C-DSC module structure;
[0038] Figure 3 Schematic diagram of the structure of the thyroid classification model based on multi-view integration.
[0039] Figure 4 Schematic diagram of the integrated network experimental results DETAILED DESCRIPTION
[0040] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation method of the present invention is described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the invention, so the present invention is not limited by the specific implementation disclosed below.
[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The present invention is further described in detail below in conjunction with the accompanying drawings and examples.
[0042] This example uses thyroid medical CT images as the identification and classification object, and performs classification and recognition according to the following method.
[0043] like Figure 1 As shown, a medical image recognition and classification method based on a multi-view integrated network includes the following steps:
[0044] Step 1: Collect thyroid CT image dataset, preprocess and manually annotate.
[0045] Step 2: Data preprocessing: crop the region of interest from the labeled data to obtain image R, and perform data enhancement (at least one of flipping, dilation, and scaling) with the region of interest as the center.
[0046] Step 3: First, enhance the feature expression through the C-DSC module.
[0047] Step 4: After pre-training (AlexNet, DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet), an integrated classification network is formed, and finally the feature map f2 is output.
[0048] Step 5: Input the processed data into the integrated network to obtain the final thyroid benign and malignant classification results.
[0049] like Figure 2 As shown, the specific process of the multi-view integrated thyroid classification model is as follows:
[0050] The uniform size of the thyroid CT image dataset acquired in step 1 is scaled to 224*224 pixels.
[0051] The image data is labeled in step 2, specifically, the acquired image is manually labeled, and a region of interest is framed to obtain an image R as a candidate label for the thyroid CT image.
[0052] The expansion of the region of interest is specifically carried out in a counterclockwise direction with the region of interest as the center, and each expansion ratio is 1 / 4 of the length and width of the original region of interest, generating N regions of interest with different expansion rates ( Figure 3 In this example, N is 4.
[0053] At this point, the original grayscale region of interest O is superimposed, with a total of N+1 (C and O) views of channels to be processed.
[0054] In step 3, the Curvelet transform is to transform the input image in random directions and angles, obtain the high-frequency and low-frequency feature maps of the image, and superimpose them with the original grayscale image to form a 3-channel feature map.
[0055] The integrated networks (AlexNet, DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet) map the output layer of each network into 2 dimensions.
[0056] The C-DSC module described in step 3 is composed of Curvelet and depthwise separable convolution.
[0057] like Figure 3 As shown, the specific working process of C-DSC is as follows:
[0058] Step 3.1: For the 5 views of the region of interest in the same batch, perform Curvelet transform according to steps: 3.1a-3.1c, feature map X.
[0059] Step 3.1a: Convert it from spatial domain to frequency domain representation via Fast Fourier Transform (FFT).
[0060] Step 3.1b: Set random directions and angles to rotate and resample the frequency domain data to capture more diverse edge and directional features. Finally, the frequency domain information is converted to the spatial domain through the inverse fast Fourier transform (IFFT).
[0061] Step 3.1c: The obtained high-frequency component H, low-frequency component L and original grayscale image D are formed into a three-channel input X.
[0062] Step 3.2: Input X into the depthwise separable convolution module and perform the depthwise separable convolution operation according to steps 3.2a-3.2b to obtain the feature map X1′.
[0063] Step 3.2a: Input X into the depthwise separable convolution. First, it undergoes a 3*3 convolution to obtain X′, whose feature map size is the same as X. After the ReLU activation function, the size remains unchanged.
[0064] Step 3.2b: After 1*1 convolution and information fusion, the feature map X1′ is obtained. After a ReLU activation function, the size of the feature map remains unchanged, which is H*W*C, and finally the feature map D is obtained; where H is the height of the image pixel, W is the width of the image pixel, and C is the number of convolution channels.
[0065] In step 4, the integrated classification network is composed of a certain number of expert subnetworks Expert Net, which is any AlexNet, DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet network structure. The number of ExpertNets is related to the number of channels of the view to be processed N+1, which is N+2. The first N+1 Expert Nets are used to process each channel view separately after the pre-training model freezes the parameters and fuses the output candidate recognition result f2. The last Expert Net is used to process the original view under non-frozen parameters and output the candidate recognition result x6. The results f2 and x6 are further fused to obtain f3 as the final classification result output. In the training process, the network parameters with the best performance on the test set are used as the network parameters of the ideal thyroid classification model.
[0066] The details are as follows:
[0067] Step 4.1: Input the feature map D into the first five Expert Nets with different freezing parameters, where ExpertNet6 is not frozen and the input is the region of interest image with the largest expansion scale.
[0068] Step 4.2: The first five Expert Net networks are used to train the views of N+1 (C and O) channels to be processed, and the output classification results are x1, x2, x3, x4, x5 respectively. The results are concatenated into feature map f1, and the Cat operation is performed to change the feature map from f1 to f2. In this example, the first five Expert Nets use AlexNet, DenseNet, GoogleNet, ResNet, and ShuffleNet with frozen parameter operation as slave networks.
[0069] Step 4.3: The result x6 output by Expert Net6 is concatenated with f2, and the Cat operation is performed to become f3, which is output as the final recognition result. In this example, Expert Net6 uses the Efficientnet backbone network.
[0070] Among them, the EfficientNet backbone network and the AlexNet, DenseNet, GoogleNet, ResNet, and ShuffleNet networks with frozen parameter operation, the last fully connected layer of each network is set to 2 units.
[0071] First, layer_fc1 is used for processing: the first connection layer processes the networks with frozen parameters: DenseNet, GoogleNet, ResNet, ShuffleNet, and Efficientnet, concatenates the final feature vectors of the five networks, and expands them to 128 feature maps after processing by the Fc fully connected layer; then after BatchNorm1d, Dropout and activation function LeakyReLU, another Fc fully connected layer operation is performed to reduce the 128 feature maps to 64 feature maps; then after the activation function LeakyReLU, the next Fc fully connected layer operation is performed to change the 64 feature maps to 10 feature maps; and then the last LeakyReLU activation operation is performed.
[0072] Secondly, layer_fc2 is used for processing: the output layer of the backbone network Efficientnet is combined with the output of layer_fc1, and expanded to 128 feature maps after processing by the first Fc fully connected layer; then after BatchNorm1d, Dropout and activation function LeakyReLU, it is reduced from 128 feature maps to 64 feature maps after another Fc fully connected layer operation; then after passing through the activation function LeakyReLU, it is converted from 64 feature maps to the score of one category corresponding to each dimension after the next Fc fully connected layer operation.
[0073] In step 5, the medical image to be identified is preprocessed, data enhanced, and the C-DSC module is used to enhance the feature expression, and then input into the thyroid ideal classification model for classification and identification, and the feature map is output.
[0074] Example:
[0075] The modeling steps of the present invention are:
[0076] Step 1: Data collection: The data of the present invention comes from thyroid CT image data collected by the First Affiliated Hospital of Shenyang Medical University.
[0077] Step 2: Data preprocessing, manual labeling of the collected data.
[0078] Step 3: Expand the marked area in a counterclockwise direction, and each expansion ratio is 1 / 4 of the length and width of the original region of interest, generating 4 regions of interest with different expansion rates.
[0079] Step 4: The preprocessed dataset is scaled to a uniform size of 224*224 pixels.
[0080] Step 5: Perform Curvelet transform and depth separable processing on the data after the data is unified in size.
[0081] Step 6: Use the pre-trained DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet, and AlexNet in ImageNet to form an integrated network. Convert the last fully connected layer of the above networks into a 2D map. Use the backbone network Efficientnet to not freeze, and freeze the other branches. Use the thyroid CT image dataset for training, and use the network parameters with the best performance in the test set as the thyroid classification model.
[0082] Steps: Resize the image or data set that needs to be classified sent by the user, and then pass it into the thyroid classification network, and finally output the thyroid classification and recognition results based on the image.
[0083] like Figure 4 As shown in the figure, the schematic diagram of the integrated network experimental results, the specific workflow is as follows:
[0084] Step 1: Compare the proposed network with ensemble networks of different numbers.
[0085] Step 2: This paper uses EfficientNet as the backbone network, and achieves the best results with 5 auxiliary networks and 6 integrated networks, with the highest accuracy index reaching 86.22 and the F1 index reaching 84.00.
[0086] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any way. Any simple modification, change and equivalent structural change made to the above embodiment based on the technical essence of the present invention still falls within the protection scope of the technical solution of the present invention.
Claims
1. A medical image recognition and classification method based on a multi-view integrated network, characterized in that: The following steps are involved: Collect and preprocess CT medical images to build a CT image dataset for training; Construct a medical image classification model with a multi-view integrated network, including: a C-DSC module based on Curvelet transform and deep separable convolution, and then connect it to a multi-expert integrated network; The CT image dataset is used to conduct supervised training on the medical image recognition classification model, and the model with the smallest classification error is selected as the ideal classification model; The unknown CT image is passed into the ideal classification model to obtain the recognition and classification results of the current image.
2. According to claim 1, a medical image recognition and classification method based on a multi-view integrated network is characterized in that: The pre-processing comprises the following steps: Modify image data to a uniform size; Manually mark the region of interest on the acquired CT image as the image label R; Perform data augmentation on the annotated regions of interest to expand the data set, and obtain multiple views C of N regions of interest of different sizes; The original grayscale view O is superimposed, and finally N+1 multi-views are obtained.
3. According to claim 2, a medical image recognition and classification method based on a multi-view integrated network is characterized in that: The data enhancement method is: The ROI area is cut out according to the marked box. Through the expansion operation, the center of the marked box is used as the reference, first expanding along the long side direction, then expanding along the wide side direction, each time expanding by 1 / 4 times, and each expansion is based on the previous step to obtain a new view.
4. According to claim 1, a medical image recognition and classification method based on a multi-view integrated network is characterized in that: In the C-DSC module, the data is subjected to Curvelet transformation to obtain two curvelet components (H, L) which are stacked with the original image to form a three-channel image. The spatial features of each channel are extracted through deep separable convolution, and then the grayscale, high-frequency and low-frequency features are effectively integrated.
5. According to the medical image recognition and classification method based on multi-view integrated network according to claim 4, the Curvelet transformation is: Convert it from the spatial domain to the frequency domain through Fast Fourier Transform (FFT); The frequency domain data is rotated and resampled by setting random directions and angles to capture more diverse edge and directional features; The frequency domain information is converted into the spatial domain through the inverse fast Fourier transform IFFT to obtain two curvelet components H and L; the high-frequency component H, the low-frequency component L and the original grayscale D image form a three-channel input X.
6. The medical image recognition and classification method based on multi-view integrated network according to claim 4, characterized in that: The depthwise separable convolution is: Input X into 3*3 convolution to get the feature map of the same size as X ′ ; Then the size remains unchanged after the ReLU activation function; After information fusion through 1*1 convolution, the feature map X1′ is obtained, which contains the weight information of each channel; after another ReLU activation function, the size of the feature map remains unchanged, and the feature map D is obtained.
7. The medical image recognition and classification method based on multi-view integrated network according to claim 1, characterized in that: The multi-expert integrated network includes a backbone network and a slave network, which is any one of DenseNet, GoogleNet, ResNet, ShuffleNet, Efficientnet, and AlexNet networks pre-trained on ImageNet.
8. A medical image recognition and classification method based on a multi-view integrated network according to claim 1 or 7, characterized in that: The multi-expert integrated network includes: Input the multi-feature map D into the first N+1 slave networks with different freezing parameters, and input the region of interest image with the largest expansion scale into Expert Net6; From the network, five Expert Net networks are used to train the views of N+1 (C and O) channels to be processed, and the output classification results are x1, x2, x3, x4, x5 respectively. The results are concatenated into feature map f1, and the Cat operation is performed to change the feature map from f1 to f2; The result x6 output by the backbone network Expert Net6 is concatenated with f2 and subjected to Cat operation to become f3, which is output as the final recognition result.
9. The medical image recognition and classification method based on multi-view integrated network according to claim 7, characterized in that: The backbone network adopts the Efficientnet network with unfrozen parameters, and the slave network adopts the AlexNet, DenseNet, GoogleNet, ResNet, and ShuffleNet networks with frozen parameters; and the last fully connected layer of each network in the integrated network is set to 2 units.
10. A medical image recognition and classification system based on a multi-view integrated network, characterized in that: It comprises a memory and a processor; the memory is used to store a computer program; the processor is used to implement a medical image recognition and classification method based on a multi-view integrated network as described in any one of claims 1 to 9 when loading and executing the computer program.
Citation Information
Patent Citations
Pulmonary nodule benign and malignant classification method and system based on multi-scale transfer learning
CN110852350A
Modeling method and target detection method and device based on attention balance feature pyramid
CN113378813A
Digestive tract multi-focus classification method and device based on multi-network ensemble learning
CN118447325A
Automated identification and classification of musculoskeletal abnormalities from medical images
WO2024254651A1