Multimodal and multiscale feature branch fusion classification method for anterior chamber images of keratitis

Through the multimodal multi-scale feature branch fusion classification method, the data quality and diversity problems in the anterior chamber image recognition classification of keratitis were solved, and more accurate image details were extracted, which improved the classification effect and model robustness.

CN116881849BActive Publication Date: 2025-05-23HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310909066.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-24
Publication Date
2025-05-23
Estimated Expiration
2043-07-24

AI Technical Summary

Technical Problem

In the anterior chamber image recognition classification of keratitis, the existing technology faces data quality problems, data diversity problems and feature extraction problems, resulting in poor classification results, especially inaccurate feature extraction of low-quality images, affecting the classification results.

Method used

The multimodal multi-scale feature branch fusion classification method is used to preprocess the anterior chamber slit lamp image, OCT image and text data of keratitis through multimodal pretreatment. The multi-scale self-attention fusion feature map generator is used to generate the fusion feature map, and feature extraction and fusion are performed through a hybrid attention network and a two-layer MLP fusion strategy, and the result prediction is finally performed through weighted summing and full connection layers.

Benefits of technology

This method can extract image details more accurately, reduce the impact of data quality on classification, improve the ability to identify and judge diverse images, enhance the generalization ability and robustness of the model, and avoid problems such as overfitting and gradient disappearance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881849B_ABST
    Figure CN116881849B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images. The method comprises: multi-modal input and pre-processing of the input data; the image data is subjected to a multi-scale self-attention fusion feature map generator to obtain fusion feature maps of two image modal data at different scales; the extracted feature maps are input to different branches for processing, and the feature maps obtained respectively are adaptively pooled to form respective feature vectors; finally, the two image data feature vectors and the vector of text data are weighted summed, and input into a fully connected layer and a ReLU activation function for calculation, and the final vector is obtained for result prediction. The present invention makes full use of medical imaging data and text data, mines and utilizes the correlation information between modalities, enhances the feature expression capability, improves the ability to recognize and judge diverse images, and comprehensively and accurately improves the classification performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of classification of anterior chamber images of keratitis, and in particular to a multi-modal multi-scale feature branch fusion classification method for classification of anterior chamber images of keratitis. Background Art

[0002] Image recognition and classification in medical images is one of the main research contents in the field of medical intelligence. Medical image recognition technology combines medical images with artificial intelligence technology, using computers to identify pathological areas instead of relying on traditional medical experts to calibrate pathological areas in medical images.

[0003] Deep learning is an algorithm in machine learning that is based on learning to represent data. It is one of the current research hotspots in the field of artificial intelligence. The use of deep learning technology has improved algorithm performance and work efficiency, and has also made it possible to implement technologies such as face recognition, autonomous driving, and intelligent translation. It is widely used in computer vision, speech recognition, natural language processing, audio recognition, and bioinformatics.

[0004] However, there are still some challenges in the recognition and classification of keratitis anterior chamber images, including data quality issues, data diversity issues, and feature extraction issues. Traditional classification methods often require manual design of features and models, which cannot adapt well to the differences between different samples and may not capture all important features. In addition, for low-quality images, feature extraction may be inaccurate, thus affecting the classification results. Summary of the invention

[0005] In order to achieve the above object, the present invention provides a multi-modal multi-scale feature branch fusion classification method for keratitis anterior chamber image classification, comprising the following steps:

[0006] Step 1), multi-modal preprocessing, wherein the modalities include keratitis anterior chamber slit lamp image, anterior chamber OCT image and text data corresponding to the OCT image.

[0007] Step 2), the keratitis anterior slit lamp image modal data and the anterior OCT image modal data obtained after preprocessing in step 1) are subjected to a multi-scale self-attention fusion feature map generator to obtain fusion feature maps of the two image modal data at different scales.

[0008] Step 3) extracts the two feature maps fused in step 2) and inputs them into different branches for processing; wherein:

[0009] The first branch uses a hybrid attention network to process two feature maps to obtain a feature map;

[0010] The second branch uses a two-layer MLP fusion strategy to process the two feature maps to obtain another feature map;

[0011] The two feature maps obtained from the two branches are adaptively pooled to form their respective feature vectors.

[0012] Step 4), the feature vector extracted in step 3) and the vector of the text data obtained in step 1) are fused again by weighted summation;

[0013] The fully connected layer and ReLU activation function are input for calculation to obtain the final vector, and then the result is predicted and the loss function is calculated at the same time, and the parameters are updated using the back propagation of the loss function.

[0014] The present invention also provides a multi-modal and multi-scale feature branch fusion classification device for anterior chamber images of keratitis, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multi-modal and multi-scale feature branch fusion classification method for anterior chamber images of keratitis is implemented.

[0015] The present invention also provides a computer-readable storage medium, which stores a computer program, and the computer program is used to execute the above-mentioned multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images.

[0016] Compared with the prior art, the present invention has the following effects:

[0017] 1) The present invention utilizes a method of multimodal input and multiscale feature fusion to fully extract image detail information and utilize features under different modalities and scales, thereby reducing the impact of data quality on classification, mining and utilizing the internal correlation of data, accurately identifying more detail information, and improving the ability to recognize and judge diverse images. Therefore, the present invention further improves the performance of computers in the recognition and classification of keratitis anterior chamber images, which is more in line with the data quality and data diversity challenges faced by current medical imaging.

[0018] 2) The present invention uses a hybrid attention graph fusion and weight mapping method to adjust and weight the feature graph, thereby enhancing the generalization and adaptability of the model. By applying these two methods to the design of the branch fusion network architecture, the robustness of the model is improved, and the use of branches avoids problems such as overfitting and gradient vanishing. At the same time, this design further enhances the relevance and importance of features, which helps to extract more discriminative feature representations; this method can make the model more robust and versatile, and can better cope with various problems in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1This is a schematic diagram of the present invention;

[0020] Figure 2 Flowchart of the hybrid attention network;

[0021] Figure 3 It is a flow chart of the two-layer MLP strategy;

[0022] Figure 4 is a flow chart of the method of the present invention;

[0023] Figure 5 Diagram of the multi-modal and multi-scale feature branch fusion classification device for anterior chamber images of keratitis. DETAILED DESCRIPTION

[0024] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the specific implementation of the present invention is described in detail below with reference to the accompanying drawings and examples.

[0025] The present invention aims to solve the defects of the prior art in the classification of anterior chamber images of keratitis, and proposes a multi-modal multi-scale feature branch fusion network for the classification of anterior chamber images of keratitis. The architecture diagram and flow chart of the present invention are respectively as follows: Figure 1 and Figure 4 As shown. Through multi-modal information input, including image and text input, the present invention can provide more comprehensive information. At the same time, the multi-scale feature branch fusion utilizes image features under different modalities and scales. The network effectively integrates the information of these different modalities and provides a more comprehensive, rich and complementary feature representation, thereby enhancing the analysis ability of keratitis anterior chamber images.

[0026] Through the fusion of multi-head self-attention mechanism, the present invention can effectively extract and integrate the local feature information of each, better extract the image detail information, and thus reduce the impact of data quality on classification. At the same time, it can also improve the recognition and judgment ability of diverse images, thereby improving the detection and positioning ability of important structures and abnormal areas in the anterior chamber images of keratitis.

[0027] By means of hybrid attention graph fusion and weight mapping, the present invention adjusts and weights the feature graph, thereby enhancing the generalization ability and adaptability of the model; in addition, through the design of the branch network, the present invention can make full use of feature information in different modes and scales, improve the robustness of the model, and further enhance the relevance and importance of the features. This helps to extract more discriminative feature representations and avoid problems such as overfitting and gradient disappearance.

[0028] An embodiment of the present invention discloses a multi-modal and multi-scale feature branch fusion classification method for anterior chamber images of keratitis, comprising the following steps:

[0029] Step 1), multi-modal input, the modalities include keratitis anterior chamber slit lamp image, anterior chamber OCT image and corresponding text data.

[0030] Preprocess the text data, including word segmentation, stop word removal, and stemming;

[0031] Preprocessing the keratitis anterior chamber slit lamp image and OCT image, including contrast enhancement and image size adjustment, thus obtaining the modality data of the two images, the keratitis anterior chamber slit lamp image and OCT image;

[0032] The preprocessed text data is converted into vector representation using the embedding layer of the Transformer model.

[0033] Step 2) The two keratitis image modality data obtained in step 1) are passed through a multi-scale self-attention fusion feature map generator to obtain fusion feature maps of the two image modality data at different scales.

[0034] Step 3), extract the two feature maps in step 2), and convert these two features Figure 1 The two feature maps are input to different branches for processing, where the first branch uses a hybrid attention network to process the two feature maps to obtain a feature map, and the second branch uses a two-layer MLP fusion strategy to process the two feature maps to obtain another feature map. The two feature maps obtained by the two branches are adaptively pooled to form their respective feature vectors.

[0035] Step 4) The feature vector extracted in step 3) and the vector of the text data obtained in step 1) are fused again by weighted summation, and input into the fully connected layer and ReLU activation function for calculation to obtain the final vector, and then the result is predicted and the loss function is calculated at the same time, and the loss function is used for back propagation to update the parameters.

[0036] In one embodiment, step 1) specifically includes:

[0037] 1) There are 1011 pairs of image-text pairs, and the input size is 256×256 of the slit lamp image x of the anterior chamber of keratitis 1 and anterior chamber OCT images of the same size x 2 , and the text data of the anterior chamber image of keratitis, for x 1 and x 2 The two images are preprocessed by contrast enhancement and resizing, and the contrast-limited adaptive histogram equalization method is used; the text is preprocessed by word segmentation, stop word removal and stemming.

[0038] 2) Input the text data text into the embedding layer of Transformer to extract its vector representation T.

[0039] In one embodiment, step 2) is specifically:

[0040] 3) Set the size of x to 256×256 1 and x 2 Input to the multi-scale self-attention fusion feature map generator, which includes the ResNet50 feature extraction module, the resizing module and the multi-head self-attention mechanism module; first extract the feature map formed by the third and fourth residual blocks through the ResNet50 residual network, and get x 1 and x 2 The two feature maps of different scales, x 1 The corresponding feature map is 1024*14*14 and 2048*7*7 x 2 The corresponding feature map is 1024*14*14 and 2048*7*7

[0041] 4) Four feature maps are obtained through the ResNet50 feature extraction module The resize module will By transposing the convolution to a larger scale, we get four 14*14-sized and Through convolution, we adjust the size to a small scale and obtain four 7*7-sized

[0042] 5) The four large-scale and four small-scale feature maps are adaptively fused into a large-scale feature map f through the multi-head self-attention mechanism module. 1 and the small-scale feature map f 2 :

[0043]

[0044]

[0045] In one embodiment, step 3) is specifically:

[0046] 6)f 1 and f 2 The first branch is to process f 1 and f 2 Input the hybrid attention network to calculate the feature map like Figure 2As shown in the figure, the hybrid attention network consists of a multi-scale hybrid attention map and a multi-channel attention pool. The multi-scale hybrid attention map extracts f 1 and f 2 The corresponding attention maps are then fused to form a mixed attention map Att, namely:

[0047]

[0048] Then the mixed attention map Att and f 2 Input into the multi-channel attention pool, f 2 Multiply the attention map in each channel of Att bit by bit, and then combine and convolve to obtain the feature map through global average pooling. Right now:

[0049]

[0050] Among them, MCApool stands for multi-channel attention pool.

[0051] 7) The second branch is as follows Figure 3 As shown, f 1 and f 2 The input is processed in a two-layer MLP fusion strategy. The first MLP is used to combine the feature maps and transform and map them into two feature maps of the same size.

[0052]

[0053] The second MLP is applied to one of the feature maps Calculate the weight map and compare the result with Multiply bit by bit to get the feature map Right now:

[0054]

[0055] 8) The two feature maps obtained and Adaptively pool GAP to form their own feature vectors z 1 and z 2 .

[0056] In one embodiment, step 4) is specifically:

[0057] 9) The feature vector z 1 and z 2 The vector representation T corresponding to the text data is weighted and summed by softmax calculation weights to obtain the fused vector z, namely:

[0058]

[0059]

[0060] t i is the i-th element of the text vector, Represents vector z 1 The jth element of Same reason.

[0061] 10) Finally, the fused feature vector z is input into the fully connected layer, and the final vector z' is obtained through the ReLU activation function. The prediction result of z' is calculated through the softmax function, and the loss function is calculated to back-propagate and update the parameters. The loss function consists of the cross entropy loss function, the attention loss function and the regularization term:

[0062] L=Lce+λ 1 ×Latt+l 2 ×Lreg

[0063] Where Lce represents the cross entropy loss function, which is used to measure the difference between the predicted value and the true label in the classification problem; Latt represents the attention loss function, which is used to measure the similarity between the mixed attention map and the feature map; Lreg represents the regularization term, which is used to control the size of the weight to prevent overfitting.

[0064] Verification example:

[0065] The test was conducted on a dataset of keratitis anterior chamber images provided by an ophthalmology hospital. Relevant practices were conducted on the three most common keratitis in clinical practice: viral keratitis, bacterial keratitis, and fungal keratitis.

[0066] This method processes the slit lamp images, OCT images and corresponding text data of the anterior chamber of keratitis, and finally obtains 1011 slit lamp images of the anterior chamber of parakeratitis, 1011 OCT images of the anterior chamber of parakeratitis and 1011 corresponding text data, forming 1011 pairs of image-text pairs. Each image and each text data corresponds to only one of the three categories of bacterial keratitis, viral keratitis and fungal keratitis.

[0067] This verification example integrates three modalities, namely slit lamp images, OCT images and corresponding text data, and distinguishes the category of keratitis anterior chamber images; in the evaluation of modality fusion, entropy, spatial frequency, mutual information, difference correlation, visual information fidelity, quality assessment index and structural similarity index are used for evaluation, and the experimental results obtained are shown in Table 1. It can be seen that the fusion effect of the present invention is good and can make full use of the information of the three modalities; in the discrimination of the category of keratitis anterior chamber images, accuracy, precision, recall rate, F1 value, specificity and false positive are used for evaluation. The experimental results obtained are shown in Table 2. It can be seen that the classification method of the present invention has a good classification effect, can preliminarily distinguish the category to which the image belongs, and provide assistance to doctors.

[0068] Table 1

[0069]

[0070] Table 2

[0071] Evaluation indicators ACC Precision Recall Test 0.943125 0.9214626 0.916213 Evaluation indicators F1 Specificity FPR Test 0.918686 0.980442 0.019557

[0072] The embodiment of the present invention also discloses a multi-modal and multi-scale feature branch fusion classification device for anterior chamber images of keratitis. Figure 5 , including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images is implemented.

[0073] The embodiment of the present invention further discloses a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is used to execute the above-mentioned multi-modal multi-scale feature branch fusion classification method for keratitis anterior chamber images. A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods.

[0074] The embodiments of the present invention described above do not constitute a limitation on the protection scope of the present invention. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.

Claims

1. Multi-modal and multi-scale feature branch fusion classification method for anterior chamber images of keratitis, Features The method comprises the following steps: Step 1), multi-modal preprocessing, wherein the modalities include keratitis anterior chamber slit lamp image, anterior chamber OCT image and text data corresponding to the OCT image; Step 2), the keratitis front slit lamp image modality data and the front OCT image modality data obtained after the preprocessing in step 1) are subjected to a multi-scale self-attention fusion feature map generator to obtain fusion feature maps of the two image modality data at different scales; Step 3) extracts the two feature maps fused in step 2) and inputs them into different branches for processing; where: The first branch uses a hybrid attention network to process two feature maps to obtain a feature map; The second branch uses a two-layer MLP fusion strategy to process the two feature maps to obtain another feature map; Adaptively pool the two feature maps obtained from the two branches to form their respective feature vectors; Step 4), the feature vector extracted in step 3) and the vector of the text data obtained in step 1) are fused again by weighted summation; The fully connected layer and ReLU activation function are input for calculation to obtain the final vector, and then the result is predicted and the loss function is calculated at the same time, and the parameters are updated using the back propagation of the loss function.

2. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 1, Features: In step 1) Preprocessing of keratitis anterior chamber slit lamp images and anterior chamber OCT images: including contrast enhancement and image resizing, Preprocess the text data: including word segmentation, stop word removal and stemming.

3. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 2, Features: The preprocessed text data is converted into vector representation using the embedding layer of the Transformer model.

4. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 1, Features: The hybrid attention network described in step 3) consists of a multi-scale mixed attention map and a multi-channel attention pool.

5. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 1, Features: The multi-scale self-attention fusion feature map generator described in step 2) includes a ResNet50 feature extraction module, a resizing module, and a multi-head self-attention mechanism module.

6. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 5, Features: The ResNet50 feature extraction module is used to obtain feature maps of two different scales of the keratitis anterior slit lamp image modality data and the anterior OCT image modality data; thereby obtaining four feature maps; The resizing module is used to resize the four feature maps to obtain four large-scale and four small-scale feature maps; The multi-head self-attention mechanism module is used to fuse four large-scale and four small-scale feature maps respectively to form two feature maps of different scales.

7. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 1, Features: The two-layer MLP fusion strategy described in step 3) is specifically: The first MLP is used to transform and map the feature maps into two feature maps of the same size after combining them; The second MLP computes a weight map for one of the feature maps and multiplies the result bit by bit with the other feature map.

8. The multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images according to claim 1, Features: In step 4), the weights of the two feature vectors and the vector corresponding to the text data are calculated through softmax.

9. Multi-modal and multi-scale feature branch fusion classification equipment for anterior chamber images of keratitis, It is characterized in that include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, and the computer program is used to execute the multi-modal and multi-scale feature branch fusion classification method for keratitis anterior chamber images described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-modal fusion classification optimization method considering inter-modal semantic distance measurement

    CN113343974A

  • Multi-modal image fusion method based on multi-scale feature extraction

    CN116071282A