Medical image classification method based on cascade fine-grained attention localization

Through the cascaded fine-grained attention positioning method, the cascaded fine-grained attention modules by using the generator and classifier, the problem of difficulty in accurately positioning fine-grained attention modules in the existing technology is solved, effectively extracting fine-grained information, and improving the accuracy of medical image classification.

CN120355975APending Publication Date: 2025-07-22BIG DATA & INFORMATION TECH RES INST OF WENZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510346515.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing medical image classification methods are difficult to accurately locate the location of subtle lesions, resulting in the inability to effectively extract fine-grained information, affecting the classification accuracy.

Method used

The cascaded fine-grained attention positioning method is adopted, and the fine-grained attention modules are cascaded by the generator and classifier to adaptively pay attention to the key areas of the image, gradually improve the positioning accuracy of the lesion position, and use the generator and classifier to train the model.

Benefits of technology

It realizes accurate positioning of subtle lesions and effective extraction of fine-grained information, improves the accuracy of medical image classification, and is especially suitable for the classification of pneumonia and chronic obstructive pulmonary diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355975A_ABST
    Figure CN120355975A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image classification method based on cascade fine-grained attention positioning, and belongs to the field of medical image analysis. A disease category corresponding to a medical image is predicted by using a generator and a classifier, the generator is used for generating feature maps of different levels of an input image, a plurality of fine-grained attention modules are cascaded behind the feature maps, and fine lesion positions are positioned through the cascaded fine-grained attention modules in training stages of the generator and the classifier to extract fine-grained information. In each level of fine-grained attention module, fusing a given first-level feature map and a given final-level feature map to obtain an attention map with the same size as the original medical image; calculating a maximum local connected region in the attention graph, and cutting the original medical image based on the maximum local connected region to obtain a fine-grained sub-graph; and adjusting the fine-grained subgraph to be the same as the original medical image in size, and repeating the calculation process of the generator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical image analysis, and particularly relates to a medical image classification method based on fine-grained attention and cascaded localization. Background Art

[0002] Fine-grained classification of medical images plays a crucial role in disease diagnosis, especially in identifying subtle pathological changes, and has been widely applied to clinical tasks such as tumor prediction. Currently, the mainstream method is to use convolutional neural networks to extract features from medical images and implement subsequent downstream tasks. Although such methods have made remarkable progress in this field, the existing methods still face the limitation of effectively obtaining sufficient local fine-grained information from global context features to improve the classification accuracy.

[0003] To address the above-mentioned medical image classification problem, researchers at home and abroad have begun to focus on fine-grained image classification methods to capture the subtle differences in diseases in medical images, thereby improving classification performance. The technical solutions relatively close to the present invention include: Y. Lu et al. (Lu Y, Yang H, Asad Z, et al. Holistic Fine-grained GGS Characterization: From Detection to Unbalanced Classification[J]. arXiv preprint arXiv:2202.00087, 2022.) used a strategy combining lesion area detection and feature enhancement. R. Dan et al. (Dan R, Li Y, Wang Y, et al. CDNet: contrastive disentangled network for fine-grained image categorization of ocular B-scan ultrasound[J]. IEEE Journal of Biomedical and Health Informatics, 2023, 27(7): 3525-3536.) utilized contrastive learning to separate key pathological features to improve classification performance. W. Park et al. (Park W, Ryu J. Fine-Grained Self-Supervised Learning with Jigsaw puzzles for medical image classification[J]. Computers in Biology and Medicine, 2024, 174: 108460.) showed potential using self-supervised pre-training methods with limited labeled data. And W. Liu et al. (Liu W, Juhas M, Zhang Y. Fine-grained breast cancer classification with bilinear convolutional neural networks(BCNNs)[J]. Frontiers in genetics, 2020, 11: 547327.) supplemented the deficiencies of traditional CNNs by mining the interaction relationships between feature channels.

[0004] In summary, among the numerous solutions for medical image classification, the performance of the model is still limited because it is impossible to accurately locate the position of subtle lesions to effectively extract fine-grained information. Summary of the Invention

[0005] In order to accurately locate the position of subtle lesions in medical images and effectively extract fine-grained information therefrom, the present invention provides a medical image classification method based on cascaded fine-grained attention localization; the fine-grained attention module adaptively focuses on the key regions in the image, and weights the feature map to effectively locate the specific position of the lesion, and then uses a cascaded manner to locate the potential lesion regions multiple times, thereby gradually improving the localization accuracy and better training the classification model.

[0006] In order to achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A medical image classification method based on cascaded fine-grained attention localization, which uses a generator and a classifier to predict the disease category corresponding to a medical image, and the generator and the classifier cascade a fine-grained attention module during the training stage to locate the position of subtle lesions to extract fine-grained information;

[0008] The training stage includes the following steps:

[0009] (1) Initialize the input image of the generator as the original medical image, and initialize the parameter i = 1;

[0010] (2) Use the generator to encode the image feature maps at different levels of the input image;

[0011] (3) Use the i-th level fine-grained attention module to fuse the first-level feature map generated in step (2) and the last-level feature map generated by the first run of the generator to obtain an attention map with the same size as the original medical image;

[0012] (4) Calculate the largest local connected region in the attention map, and crop the original medical image based on the largest local connected region to obtain a fine-grained sub-image;

[0013] (5) Adjust the fine-grained sub-image obtained in step (4) to the same size as the original medical image, which is used as the input image of the generator, update the parameter i to be equal to i + 1, and return to step (2) until the updated i > N, where N is the preset number of cascade layers;

[0014] (6) Input the last-level feature map generated by each run of the generator into the classifier, take the average of the classification confidence corresponding to each last-level feature map as the final prediction result, and iteratively train the generator, the classifier and the cascaded fine-grained attention modules at all levels according to the loss between the prediction result and the true label of the original medical image.

[0015] Further, the medical image is an X-ray image or a nuclear magnetic resonance image.

[0016] Further, the generator is a Resnet network structure.

[0017] Further, the calculation process of the fine-grained attention module includes:

[0018] Given the first-level feature map and the last-level feature map of the input;

[0019] Preprocess the two feature maps and calculate the outer product;

[0020] Use the fully connected layer to linearly transform the outer product result and calculate the attention map through the activation function.

[0021] Further, the preprocessing of the two feature maps and the calculation of the outer product include:

[0022] Perform convolution and pooling operations on the first-level feature map in sequence to obtain the first result;

[0023] Perform convolution operation on the last-level feature map to obtain the second result;

[0024] Use the second result as the anchor point to perform per-position interaction calculation with the first result.

[0025] Further, the dimension of the first result is c×1×1, and the dimension of the second result is c×h×w, where h and w are the height and width of the feature.

[0026] Further, the 8-connectivity algorithm is used to calculate the maximum local connected region for the attention map.

[0027] Further, N = 2.

[0028] In a second aspect, the present invention proposes a medical image classification system based on cascaded fine-grained attention localization for implementing the above-mentioned medical image classification method based on cascaded fine-grained attention localization.

[0029] Beneficial effects of the present invention:

[0030] The present invention uses the fine-grained attention and cascaded localization method to achieve fine-grained classification of medical images, especially suitable for the classification of complex diseases such as pneumonia and chronic obstructive pulmonary disease. This method obtains the position information related to accurate disease prediction through the cascaded localization strategy, extracts the corresponding fine-grained features according to this position information, so as to improve the accuracy of fine-grained classification. Specifically, the present invention uses the fine-grained attention module to realize the fine-grained interaction of features through vector outer product, generates an attention map containing lesion position information, so as to accurately locate the lesion and capture its local details. At the same time, the fine-grained attention module locates and screens the potential lesion areas in a cascaded manner multiple times, gradually improving the localization accuracy and feature extraction ability. Description of the Drawings

[0031] Figure 1Schematic diagram of the architecture of the medical image classification method based on cascaded fine-grained attention localization proposed by the present invention;

[0032] Figure 2 For the first cascaded image and attention image;

[0033] Figure 3 For the second cascaded image and attention image. Detailed implementation manners

[0034] The present invention will be further described and explained below in conjunction with the detailed implementation manners. The embodiments are only examples of the present disclosure and do not delimit the scope of limitation. The technical features of each embodiment of the present invention can be combined correspondingly without conflict.

[0035] The accompanying drawings are only schematic diagrams of the present invention and are not necessarily drawn to scale. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0036] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all steps. For example, some steps can be decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may be changed according to the actual situation.

[0037] A medical image classification method based on cascaded fine-grained attention localization proposed by the present invention is introduced respectively from three parts: model structure, training method, and actual application inference.

[0038] (1) Model structure

[0039] The model for medical image classification mainly includes a generator G and a classifier C f two parts; in this embodiment, D = {X S , Y S} is designed to represent the labeled training data set for the training stage, and D t = {X t} represents the target data set without labels for the test stage, where X ∈ {X s , X t} represents the input medical image, and Y S represents the disease category corresponding to the medical image X S .

[0040] The input of the generator G is the medical image data X, and the output result is the feature maps at different stages α represents the number of stages; the classifier Cf The input is the fused feature F f , where the fused feature F f refers to the feature map of the last layer generated by running the generator multiple times during the training phase. During the inference phase, only one layer of the feature map is required. The output of the generator G is the classification result Among them, N represents the number of disease categories corresponding to the medical image. In this example, α = 4

[0041] The network structure of the generator is a pre-given network backbone structure; in this example, the backbone structure of the network is Resnet50, α = 4, which can generate four layers of image feature maps. After this backbone structure, several groups of fine-grained attention modules are cascaded through a cascading mechanism to locate the subtle lesion positions to extract fine-grained information

[0042] In each level of the fine-grained attention module, its input is the first-stage feature map X1 with edge texture information and the last feature map X4 with rich global semantic information. Convolution processing is performed on X4 to obtain Convolution and pooling processing are performed on X1 to obtain Taking as the anchor point and performing per-position interaction calculation to obtain the interaction attention map A

[0043]

[0044] Among them, fc represents a fully connected layer, σ represents the sigmoid function represents the outer product of vectors

[0045] The 8-connectivity algorithm is used for the attention map A to obtain the largest local connected region. Based on the processed attention map A, the original image X is cropped to obtain a finer-grained sub-image I of the next cascade layer, and this sub-image is Resized to the same size as the original image

[0046] (2) Training method

[0047] Referring to Figure 1 , the training phase mainly includes the following steps

[0048] Step 1: Randomly sample n groups of training samples X from the training dataset

[0049] Step 2: Input the training data X into the generator G to obtain the global feature maps at different stages of the image

[0050] Step 3: Regard the first-layer feature map and the last-layer feature map obtained by the generator G as the global texture features and the global semantic features Two feature maps are input into the first-level fine-grained attention module. Convolution and pooling operations are sequentially performed on the first-level feature map to obtain a first result with a dimension of c×1×1. Convolution operation is performed on the last-level feature map to obtain a second result with a dimension of c×h×w, where h and w are the height and width of the feature. The outer product of vectors is calculated for the two results, and the outer product result is linearly transformed using a fully connected layer and calculated through an activation function to obtain the attention map A at layer 0 0 ;

[0051] Step 4: For the attention map A 0 The 8-connectivity algorithm is used to calculate the largest local connected region, and the original medical image is cropped based on the largest local connected region to obtain the fine-grained subgraph I 0 ;

[0052] Step 5: After Resizing the subgraph I 0 to be the same size as the original image, it is input into the same generator G to obtain the local feature map

[0053] Step 6: Using the same method, the local texture feature and the global semantic feature are input into the second-level fine-grained attention module to obtain the attention map A at layer 1 1 ;

[0054] Step 7: For the attention map A 1 The 8-connectivity algorithm is used to calculate the largest local connected region, and the original medical image is cropped based on the largest local connected region to obtain the fine-grained subgraph I 1 ;

[0055] Step 8: Repeat steps 5 - 7 for a total of λ - 1 times, where λ is the number of cascading times. In this example, λ = 2. Each repetition process is implemented using an independent fine-grained attention module

[0056] Step 9: The above-obtained λ + 1 local features are respectively input into the classifier C f to obtain all prediction confidences P i . For P i , an averaging operation is performed to obtain the final confidence P:

[0057] P = mean(P i )

[0058] Step 10: Calculate the loss function L for the final classification confidence P:

[0059] L = L ce (P, y)

[0060] where Lce represents the cross - entropy loss, and y represents the corresponding labels of the training set images;

[0061] Step 11: Generator G and classifier C f Calculate the forward - propagation error value according to the loss function L, and then perform back - propagation according to the error;

[0062] Step 12: Repeat Steps 1 to 11 until e iterations are completed, where e is the pre - given number of training rounds; in this example, e = 200.

[0063] (III) Practical application and inference

[0064] Step 1: Read a single medical image I;

[0065] Step 2: Input the image I into the generator model G and pass it through the classifier C f Calculate the prediction result

[0066] To explore the impact of different cascade layers on performance, an ablation study on the cascade layers was conducted in this embodiment. The experimental results are shown in Table 1. It can be observed that using two cascade layers can achieve the best performance, with an accuracy of 94.39%. Fewer or too many cascade layers will have a negative impact on performance. This is because too few cascade layers prevent the model from accurately locating the disease area, making it difficult to capture effective fine - grained information. On the contrary, too many cascade layers may introduce too many parameters, resulting in overfitting.

[0067] Table 1

[0068]

[0069] To prove the effectiveness of the method proposed in the present invention in locating the lesion area, the attention maps and cropped regions generated on different cascade layers of the network were visualized. The visualization results are as Figure 2 Figure 3 shown. As observed, the attention maps from earlier cascade layers tend to capture coarse - grained features, such as general anatomical structures, while the attention maps from deeper cascade layers gradually refine the localization of the lesion area with higher accuracy. This visualization result verifies the motivation behind the method of the present invention.

[0070] In this embodiment, a medical image classification system based on cascade fine - grained attention localization for implementing the above - mentioned detection method is also provided, including:

[0071] A generator for generating image feature maps of different levels of the input image;

[0072] A classifier for predicting the disease category corresponding to the medical image;

[0073] A positioning module, which includes a multi-level fine-grained attention module for locating subtle lesion positions to extract fine-grained information during the training phase;

[0074] A training module for jointly training with the positioning module, generator, and classifier, and the training process adopts the steps of the above-mentioned training phase.

[0075] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. For example, the positioning module may further include:

[0076] A fine-grained attention module for fusing a given first-level feature map and the last-level feature map to obtain an attention map with the same size as the original medical image;

[0077] A fine-grained sub-image generation module for calculating the largest local connected region in the attention map, cropping the original medical image based on the largest local connected region to obtain a fine-grained sub-image, and adjusting it to the same size as the original medical image.

[0078] The implementation methods of the remaining modules will not be elaborated here. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present invention. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0079] The embodiments of the system of the present invention can be applied to any device with data processing capabilities, and the device with data processing capabilities can be a device or apparatus such as a computer. The system embodiment can be implemented by software, or by hardware or a combination of software and hardware. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of any device with data processing capabilities reading the corresponding computer program instructions in the non-volatile memory into the memory for operation.

[0080] The above embodiments only represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the present invention. For those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention.

Claims

1. A medical image classification method based on cascaded fine-grained attention localization, characterized in that, Predict the disease category corresponding to a medical image using a generator and a classifier. The generator and the classifier cascade fine-grained attention modules during the training phase to locate subtle lesion positions and extract fine-grained information. The training phase includes the following steps: (1) Initialize the input image of the generator as the original medical image, and initialize the parameter i = 1. (2) Use the generator to encode the image feature maps at different levels of the input image. (3) Use the i-th level fine-grained attention module to fuse the first-level feature map generated in step (2) and the last-level feature map generated by the first run of the generator to obtain an attention map with the same size as the original medical image. (4) Calculate the largest local connected region in the attention map, and crop the original medical image based on the largest local connected region to obtain a fine-grained sub-image. (5) Resize the fine-grained sub-image obtained in step (4) to the same size as the original medical image, which serves as the input image of the generator. Update the parameter i to be equal to i + 1, and return to step (2) until the updated i > N, where N is the preset number of cascade layers. (6) Input the last-level feature map generated by each run of the generator into the classifier, take the average of the classification confidence levels corresponding to each last-level feature map as the final prediction result, and iteratively train the generator, the classifier, and the cascaded fine-grained attention modules at all levels according to the loss between the prediction result and the true label of the original medical image.

2. The medical image classification method based on cascaded fine-grained attention localization according to claim 1, wherein The medical image is an X-ray image or a nuclear magnetic resonance image.

3. The medical image classification method based on cascaded fine-grained attention localization according to claim 1, wherein The generator is a Resnet network structure.

4. The medical image classification method based on cascade fine-grained attention localization according to claim 1, characterized in that The calculation process of the fine-grained attention module includes: Given the input first-level feature map and last-level feature map. Preprocess the two feature maps and calculate the outer product. Use a fully connected layer to linearly transform the outer product result and calculate the attention map through an activation function.

5. The medical image classification method based on cascade fine-grained attention localization according to claim 4, wherein The preprocessing of the two feature maps and calculating the outer product includes: Perform convolution and pooling operations on the first-level feature map in sequence to obtain a first result. Perform a convolution operation on the last-level feature map to obtain a second result. Use the second result as an anchor to perform a position-by-position interaction calculation with the first result.

6. The medical image classification method based on cascade fine-grained attention localization according to claim 5, characterized in that, The dimension of the first result is c×1×1, and the dimension of the second result is c×h×w, where h and w are the height and width of the feature.

7. The medical image classification method based on cascaded fine-grained attention localization according to claim 1, characterized in that Use an 8-connectivity algorithm on the attention map to calculate the largest local connected region.

8. The medical image classification method based on cascaded fine-grained attention localization according to claim 1, wherein N=2。 9. A medical image classification system based on cascaded fine-grained attention localization, characterized in that, Includes: A generator, which is used to generate image feature maps at different levels of the input image. A classifier, which is used to predict the disease category corresponding to the medical image. A localization module, which includes multiple levels of fine-grained attention modules and is used to locate subtle lesion positions to extract fine-grained information during the training phase. A training module, which is used to jointly train the localization module, the generator, and the classifier, and the training process adopts the steps of the training phase described in claim 1.

10. The medical image classification system based on cascaded fine-grained attention localization according to claim 9, wherein, The localization module includes: A fine-grained attention module, which is used to fuse the given first-level feature map and last-level feature map to obtain an attention map with the same size as the original medical image. A fine-grained subgraph generation module, which is used to calculate the largest local connected region in the attention map, crop the original medical image based on the largest local connected region to obtain a fine-grained subgraph, and adjust it to the same size as the original medical image.