Method for constructing a fine fracture image recognition network based on a cross-attention mechanism

By constructing a fine fracture image recognition network based on cross attention mechanism, the problem of poor fracture image recognition effect in the prior art is solved, and high-precision and high-generalized fracture image recognition is achieved, which simplifies the recognition process.

CN114463312BActive Publication Date: 2025-07-25XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210126442.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-10
Publication Date
2025-07-25
Estimated Expiration
2042-02-10

AI Technical Summary

Technical Problem

In the field of medical imaging, especially in the fine recognition task of fracture imaging, it is difficult to effectively identify images in different domains and different recognition tasks, and there are problems such as insufficient feature discrimination and poor classification effect.

Method used

A fine identification network for fracture images based on the cross attention mechanism is constructed, and feature adaptation modules are combined with feature extraction modules, feature adaptation modules and image classification modules, and a cross attention network structure is adopted, and the model is trained through local and global classifier loss to improve feature adaptability and recognition accuracy.

Benefits of technology

The accuracy of fine-type recognition of fracture images skeleton parts is improved, the generalization ability and application of the method are enhanced, and the recognition process is simplified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114463312B_ABST
    Figure CN114463312B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for constructing a fine fracture image recognition network based on a cross-attention mechanism, which mainly relates to the field of medical image processing. The method includes the following steps: obtaining a sample data set of original fracture images, manually classifying and labeling, and cross-data augmentation; constructing a cross-attention network structure including a feature extraction module, a feature adaptation module, and an image classification module; forming a total loss model, and training the cross-attention network structure. The beneficial effect of the present invention is that it can solve the technical problem that the existing few-shot image recognition algorithms cannot effectively recognize images due to different domains and different recognition tasks in actual use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image processing, and specifically to a method for constructing a fine fracture image recognition network based on a cross-attention mechanism. Background Art

[0002] Multi-source image few-shot recognition algorithms have broad application prospects in the medical field. The purpose of few-shot classification is to recognize unlabeled samples when there are only a few labeled samples in the unseen domain. In the specific implementation process, few-shot learning algorithms face the following three major challenges: extremely few available samples, large differences among samples of the same category, and similarities among samples of different categories; especially in the aspect of fine image classification, the characteristic of extremely high similarity among different category samples in the unseen domain is particularly prominent. Existing few-shot recognition methods extract features from both the support set labeled samples and the validation set unlabeled samples, resulting in insufficient discriminability of features and poor classification effects.

[0003] Therefore, in terms of algorithm selection, on the one hand, a more flexible algorithm needs to be selected to adapt to the few-shot recognition problem under specific domain transfer conditions; on the other hand, the algorithm should have a higher recognition accuracy for the specific task of recognizing fine parts of fracture images. Currently, common image recognition algorithms usually cannot have both of the above characteristics. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for constructing a fine fracture image recognition network based on a cross-attention mechanism, which can solve the technical problem that existing few-shot image recognition algorithms cannot effectively recognize images due to different domains and different recognition tasks.

[0005] To achieve the above object, the present invention is realized through the following technical solutions:

[0006] A method for constructing a fine fracture image recognition network based on a cross-attention mechanism includes the following steps:

[0007] S1, obtain a sample data set of original fracture images, perform fine classification and annotation on this data set manually, and perform cross-data augmentation based on the classified sample data set;

[0008] The sample data set includes fracture image data of 17 different types of bone regions;

[0009] The fine classification and annotation of this data set includes classifying and annotating each fracture lesion site one by one, using the above fracture image data as a test data set, selecting 5 categories as a training set, randomly selecting the remaining 5 categories as a support set, and the remaining 7 categories as a validation set;

[0010] In the training stage, a mixed dataset consisting of the Mini ImageNet dataset and the datasets CUB200, Oxford Flower, and Stanford Car is selected as the training set. After removing duplicate categories, the total number of categories is 88. Additionally, 5 categories from the above fracture imaging data are used as the training set. Then, a cross-set erasing and completion module is used to finely classify the bone parts of the fracture images, and a small number of image samples in the original training set are subjected to cross-data augmentation. The cross-data augmentation includes:

[0011] S1.1, through the erasing module, erase the pixels in the support set image data that contribute highly to the final classification result;

[0012] S1.2, through the pixel completion model, complete the erased pixel points to generate the completed image data;

[0013] S1.3, add the generated data to the validation set of the fracture image samples of the same task for data cross-fusion;

[0014] S1.4, the validation set samples and the support set samples are used to test the few-shot image recognition network model for pre-training;

[0015] In the testing stage, in each trial of the model, 5 categories are randomly selected from the 12 categories after removing the training set as the validation set, and the remaining 7 categories are used as the test set to ensure the high independence of each part of the data and the high effectiveness of the recognition result;

[0016] Then, through the enhanced image data with annotations, before performing the few-shot image recognition task through the pre-trained cross-attention network, the training set images are statistically enhanced again, including random flipping, distortion, expansion, and cropping of the images.

[0017] S2, construct a cross-attention network structure including a feature extraction module, a feature adaptation module, and an image classification module;

[0018] S3, form a total loss model from the local classifier loss and the global classifier loss to train the cross-attention network structure.

[0019] The feature distribution map obtained based on the feature extraction module is input into the feature adaptation module. Through metric learning, appropriate feature representations are obtained for each pair of support classes and validation samples. The feature adaptation module takes the cross-attention module as the core and also includes a feature correlation module and a meta-fusion module.

[0020] The image classification module consists of a local classifier and a global classifier in parallel. Among them, the local classifier is used to calculate the cosine distance between the support set features and the validation set features, calculate the similarity between the two features, and thus obtain the probability value of the validation set features; the global classifier passes through a fully connected layer and then performs classification through Softmax.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0022] The designed cross-attention few-shot image recognition network improves the accuracy of fine classification of the bone parts in fracture images. At the same time, the method is simple, and has strong generalization and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 It is a flowchart of a method for constructing a fine fracture image recognition network based on a cross-attention mechanism according to an embodiment of the present invention;

[0024] Figure 2 It is a structural framework diagram of a fine fracture image recognition network based on a cross-attention mechanism according to an embodiment of the present invention;

[0025] Figure 3 It is a schematic diagram of the meta-fusion module according to an embodiment of the present invention;

[0026] Figure 4 It is a schematic diagram of the erasing and complementing module according to an embodiment of the present invention;

[0027] Figure 5 It is the result of the ablation experiment comparison experiment;

[0028] Figure 6 It is the result of the comparison experiment of different data sets. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. And modifications made based on the present invention, these equivalent forms also fall within the scope defined by this application.

[0030] As Figure 1 shown, constructing a fine fracture image recognition network based on a cross-attention mechanism includes the following steps:

[0031] Step 1: Make a few-shot image data set of original fracture images, perform fine classification on multi-site fracture images, and perform cross-data augmentation on the original small amount of images;

[0032] Step 2: Construct a cross-attention network structure, including a feature extraction module, a feature adaptation module, and an image classification module;

[0033] Step 3: Based on the few-shot recognition network, the total loss consists of the local classifier loss and the global classifier loss, and train the cross-attention network structure.

[0034] First, as Figure 1 shown, make a few-shot image dataset of original fracture images, perform fine classification on whole-body multi-site fracture images, and perform cross-data augmentation on the original small number of images, which specifically includes:

[0035] Taking the Department of Orthopedics of Peking Union Medical College Hospital as the main data collection point and the township hospitals as the auxiliary data collection points, collect whole-body multi-site fracture image data (mainly X-ray film data), obtain X-ray film image data of fractures in 17 different types of bone regions, a total of 206 X-ray film images as the training dataset, and classify and label the fracture lesion sites by a professional orthopedic doctor team; specifically divided into 17 fine categories: humerus, scapula, clavicle, rib, vertebra, pelvis, femur, patella, tibia, fibula, calcaneus, metacarpal, carpal, ulna and radius, and phalanges.

[0036] Use 56 X-ray film image data of fractures in 17 different types of bone regions collected as the test dataset. The dataset contains five main bone region categories of the human body, namely the skull, upper trunk, lower trunk, lower limbs and upper limbs. The data is subdivided into 17 fine-grained categories, including the skull, clavicle, calcaneus (lower limb), metacarpal (upper limb), tibia, fibula, etc. Select 5 categories as the training set, randomly select 5 categories from the remaining 12 categories as the support set, and the remaining 7 categories as the validation set.

[0037] In the training stage, select a mixed dataset composed of the Mini ImageNet dataset and the datasets CUB200, Oxford Flower, and Stanford Car as the training set. After removing the duplicate categories, the total number of categories is 88, and add 5 categories from the above-mentioned 17 fine-class fracture images manually divided as the training set.

[0038] Next, as Figure 4 shown, adopt the cross-set erasing-completing module to perform fine classification on the bone parts of fracture images and perform cross-data augmentation on a small number of image samples in the original training set. The cross-set erasing-completing module mainly consists of two parts: erasing-completing and cross-set data augmentation. First, use the erasing-completing module to process the validation set data to generate new validation set images, and then use the cross-set data augmentation module to construct multiple new tasks. Finally, use the new tasks to train the model. The cross-set erasing-completing module is only used in the training stage, so it will not increase the model inference time.

[0039] In the test stage, for each trial of the model, 5 out of the 12 categories (1 sample per category, i.e., 5-way-1-shot) are randomly selected from the remaining data after removing the training set as the validation set, and the remaining 7 categories are used as the test set to ensure the high independence of each part of the data and the high effectiveness of the recognition results.

[0040] Next, the labeled image data after the training data is enhanced by data augmentation passes through the pre-trained cross-attention network. Before performing the few-shot image recognition task, the training set images are statistically enhanced again, including random flipping, distortion, expansion, and cropping of the images. Specifically, it includes: 1) Random scaling, normalizing the image size to between -0.5 and 0.5; 2) Randomly flipping the image left and right and randomly distorting it; 3) Randomly cropping the image, with the aspect ratio of the cropped area being 0.5 to 2, and the effective IOU cropping thresholds being 0, 0.1, 0.3, 0.5, 0.7, 0.9, and the ratio of the cropped area to the original image being 0.3 to 1.

[0041] As Figure 2 shown, the fine fracture image recognition network model of the cross-attention mechanism also includes a ResNet-12 residual learning structure as the backbone structure part of the feature extraction module. The input image size is 84×84. During the training process, the optimizer is selected as SGD.

[0042] For the feature extraction module, the present invention adjusts the feature extraction module from two aspects. On the one hand, since each channel of the feature map represents a different feature distribution pattern, to highlight the important feature patterns, the model learns the channel attention of the support set and the validation set; on the other hand, to highlight the target object, the model of the present invention only learns the spatial attention of the support set, and the spatial attention feature map is generated by the support set samples to improve the adaptability of the features to the current task.

[0043] Next, the feature distribution map after feature extraction is input into the feature adaptation module, and appropriate feature representations are obtained for each pair of support classes and validation samples through metric learning. The core of the feature adaptation module is the CrossAttention Module (CAM), and this module also includes a feature correlation module and a meta-fusion module.

[0044] First, the correlation graph between the support set feature map P and the validation set feature map Q is calculated through the feature correlation module, and then the correlation graph is used to guide the generation of the cross-attention graph. The feature correlation module calculates the semantic correlation between the spatial positions of P and Q using the cosine distance to obtain the correlation relationship graph R.

[0045] Based on the correlation relationship graph R, two correlation mappings are defined: the support set correlation mapping R p and the validation set correlation mapping R q, where represents the correlation between the local support set eigenvector and all validation set eigenvectors, represents the correlation between the local validation set eigenvector and all support set eigenvectors. In this way, R p and R q characterize the local correlation between the class and the validation set feature mapping.

[0046] Next, as Figure 3 shown, the correlation graph R output by the feature correlation module is input into the meta-fusion module, and the meta-fusion layer generates attention maps for the support set and the validation set respectively according to the corresponding ones. The meta-fusion layer takes the correlation map R p / R q as the input, and uses kernel to perform convolution operations. For each pair of class and validation features, each local correlation vector of R p / R q is fused into an attention scalar to aggregate the correlation between features. Then, the Softmax function is used to normalize the attention scalar to obtain the final classification result.

[0047] Furthermore, the loss function is calculated through the image classification module, and the few-shot image recognition network based on the cross-attention mechanism is trained by minimizing the classification loss of the training set query samples. The classification module consists of a local classifier and a global classifier.

[0048] The following local classification loss is adopted:

[0049] ;

[0050] where is the probability of predicting the th class as . At the same time, the global classification loss L1 is constructed, and the global classification loss is expressed as:

[0051] ;

[0052] Finally, the overall classification loss is defined as ;

[0053] where λ is the weight for balancing the loss . The gradient descent algorithm is used to optimize , and the network is trained end-to-end.

[0054] The advantage of the present invention lies in the designed cross-attention few-shot image recognition network, which improves the accuracy of fine-class recognition of the bone part in fracture images. At the same time, the method is simple, and has strong generalization and applicability.

Claims

1. A method for constructing a fine recognition network for fracture images based on a cross-attention mechanism, characterized in that It includes the following steps: S1. Obtain the sample data set of the original fracture images, finely classify and label the data set manually, and perform cross-data augmentation based on the classified sample data set; The sample data set includes fracture image data of 17 different types of bone regions; The fine classification and labeling of the data set includes classifying and labeling each fracture lesion site one by one, using the above fracture image data as the test data set, selecting 5 categories as the training set, randomly selecting the remaining 5 categories as the support set, and the remaining 7 categories as the validation set; In the training stage, select the mixed data set composed of the Mini ImageNet data set, the data set CUB200, Oxford Flower, and StanfordCar as the training set. After removing the duplicate categories, the total number of categories is 88. Add 5 categories from the above fracture image data as the training set. Then, use the cross-set erasure-completion module to finely classify the bone parts of the fracture images and perform cross-data augmentation on a small number of image samples in the original training set. The cross-data augmentation includes: S1.

1. Through the erasure module, erase the pixels with high contribution to the final classification result in the support set image data; S1.

2. Through the pixel completion model, complete the erased pixel points to generate the completed image data; S1.

3. Add the generated data to the fracture image sample validation set of the same task for data cross-fusion; S1.

4. The validation set samples and the support set samples are used to test the few-shot image recognition network model for pre-training; In the testing stage, in each trial of the model, randomly select 5 categories from the 12 categories after removing the training set as the validation set, and the remaining 7 categories as the test set to ensure the high independence of each part of the data and the high effectiveness of the recognition result; Then, through the augmented data of the annotated images, before performing the few-shot image recognition task through the pre-trained cross-attention network, perform statistical data augmentation on the training set images again, including random flipping, distortion, expansion, and cropping of the images; S2. Construct a cross-attention network structure including a feature extraction module, a feature adaptation module, and an image classification module; S3. Build a total loss model from the local classifier loss and the global classifier loss, and train the cross-attention network structure.

2. The method for constructing a fine fracture image recognition network based on the cross-attention mechanism according to claim 1, wherein The feature distribution map obtained based on the feature extraction module is input into the feature adaptation module. Through metric learning, appropriate feature representations are obtained for each pair of support classes and validation samples. The feature adaptation module takes the cross-attention module as the core and also includes a feature association module and a meta-fusion module.

3. The method for constructing a fine fracture image recognition network based on the cross-attention mechanism according to claim 2, wherein, The image classification module consists of a parallel local classifier and a global classifier. Among them, the local classifier is used to calculate the cosine distance between the support set features and the validation set features, calculate the similarity between the two features to obtain the probability value of the validation set features; the global classifier passes through a fully connected layer and then performs classification through Softmax.