Image processing method and system based on multiple attention, and terminal
By introducing a two-stage attention fusion mechanism in image processing, combining multi-head self-attention and convolutional block attention modules, effectively fusing global and local features, the problem of insufficient feature fusion in the existing technology is solved, and the accuracy of image classification and the robustness of the model are improved.
Patent Information
- Application Number
- CN202411992400.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-13
AI Technical Summary
When combining the multi-head attention mechanism with convolutional neural network for image processing, the prior art focuses on the use of multi-head attention, ignoring the further extraction of local features and the effective fusion of the two types of features, resulting in inaccurate image classification results.
An image processing method based on the dual-stage attention fusion mechanism is proposed. By constructing a feature extraction module containing multi-head self-attention and convolutional block attention modules, combining global and local features, and performing feature fusion through a two-step attention fusion unit, the image processing model is finally constructed using the ResNet architecture.
It significantly improves the accuracy of image processing, enhances the model's detection ability of complex lesion areas, dynamically adjusts the weight allocation of global and local features, and improves the robustness and adaptability of the model.
Smart Images

Figure CN119992165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and in particular to a multi-attention based image processing method, system, terminal and computer-readable storage medium. Background Art
[0002] In the field of image analysis, image classification is a key task that involves categorizing images according to specific structures. The main goal is to automatically identify and classify these images based on their content using computer algorithms. Image classification generally involves two key steps: extracting effective features from images and using these features to build a model for classifying image datasets. Current methods rely heavily on machine learning or deep learning models.
[0003] Although these methods show great potential for intelligent classification, they also face some challenges. A major challenge is to capture the characteristics of images, which show high similarity and complex details. A large number of studies have shown that convolutional neural network and Transformer based models have achieved excellent performance in image classification tasks. Although convolutional neural networks are good at extracting local features, they often have difficulty in capturing the global context of the image, which is crucial for accurate classification.
[0004] Recognizing the complementary advantages of convolutional neural networks and transformers, a series of hybrid architectures have been proposed to integrate convolutional neural networks with transformers to more effectively address various challenges in medical imaging. Although these hybrid models perform well in extracting local and global features, they usually require large training datasets to succeed. Building such large datasets requires a lot of time and resources. In addition, the quadratic computational complexity of transformers leads to high computational costs, posing a major challenge to resource-constrained clinical applications. To address this issue, it was found that the multi-head attention mechanism alone can extract global features. Although the combination of multi-head attention mechanism with convolutional neural networks can effectively extract global and local features, these studies often focus on the use of multi-head attention and ignore the further extraction of local features and the effective fusion of the two types of features.
[0005] Therefore, the prior art still needs to be improved and developed. Summary of the invention
[0006] The main purpose of the present invention is to provide an image processing method, system, terminal and computer-readable storage medium based on multi-attention, aiming to solve the problem that when the multi-head attention mechanism is combined with a convolutional neural network for image processing in the prior art, the use of multi-head attention is often focused on, while the further extraction of local features and the effective fusion of the two types of features are ignored, resulting in the problem that the image classification results are not accurate enough.
[0007] To achieve the above object, the present invention provides an image processing method based on multi-attention, and the image processing method based on multi-attention comprises the following steps:
[0008] Acquire an image to be processed, and preprocess the image to be processed to obtain a target image;
[0009] Constructing an image processing model based on a dual-stage attention fusion mechanism, training and testing the image processing model to obtain a target model;
[0010] The target image is input into the target model for feature extraction and classification to obtain a classification result of the image to be processed.
[0011] Optionally, in the multi-attention based image processing method, the target model comprises: a first scale feature extraction module, a second scale feature extraction module, a third scale feature extraction module, a fourth scale feature extraction module, a global average pooling layer and a classifier;
[0012] Among them, the first scale feature extraction module includes a convolution layer and an MFF Block submodule, and the second scale feature extraction module, the third scale feature extraction module and the fourth scale feature extraction module all include an average pooling layer and an MFF Block submodule.
[0013] Optionally, in the multi-attention based image processing method, the MFF Block submodule comprises: a global feature path, a local feature path and a two-step attention fusion unit;
[0014] The step of inputting the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed specifically includes:
[0015] Inputting the target image into the convolution layer of the first scale feature extraction module for convolution to obtain a first input feature, inputting the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first global feature, and inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first local feature;
[0016] Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature;
[0017] Inputting the first output feature to the average pooling layer of the second scale feature extraction module for average pooling to obtain a second input feature, inputting the second input feature to the global feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second global feature, and inputting the first output feature to the local feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second local feature;
[0018] Inputting the second global feature and the second local feature into the two-step attention fusion unit of the MFF Block submodule of the second scale feature extraction module for feature fusion to obtain a second output feature, and continuously inputting the second global feature into the next scale feature extraction module for feature extraction and fusion, until the target output feature is obtained after passing through the fourth scale feature extraction module;
[0019] The target output feature is input into the global average pooling layer for dimensionality reduction to obtain a vector for classification, and after the vector passes through the classifier, the classification result of the image to be processed is obtained.
[0020] Optionally, the multi-attention based image processing method, wherein the step of inputting the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first global feature, specifically includes:
[0021] Input the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module to perform an all2all attention calculation, and linearly transform the first input feature with a query weight matrix, a keyword weight matrix, and a position encoding weight matrix, respectively, to obtain a query vector, a key vector, and a value vector;
[0022] Encode the height and width of the first input feature to obtain a position encoding matrix, multiply the position encoding matrix by the query vector to obtain a first intermediate vector, and multiply the key vector by the query vector to obtain a second intermediate vector;
[0023] The first intermediate vector and the second intermediate vector are summed element by element, the sum result is input into the softmax layer for normalization to obtain a normalized result, and the normalized result is multiplied by the value vector to obtain a first global feature.
[0024] Optionally, the multi-attention based image processing method, wherein the step of inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first local feature, specifically includes:
[0025] Inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module to perform a maximum pooling operation and an average pooling operation to obtain a first maximum pooling result and a first average pooling result;
[0026] Inputting the first maximum pooling result and the first average pooling result into a shared multi-layer perceptron for feature extraction to obtain a first intermediate feature and a second intermediate feature;
[0027] The first intermediate feature and the second intermediate feature are summed element by element, and the summation result is passed through a sigmoid layer to obtain channel attention;
[0028] Performing a maximum pooling operation and an average pooling operation on the channel attention to obtain a second maximum pooling result and a second average pooling result;
[0029] A feature expression of the target local area is extracted from the second maximum pooling result and the second average pooling result, and the feature expression is passed through a sigmoid layer to obtain spatial attention, and the spatial attention is used as the first local feature.
[0030] Optionally, the multi-attention based image processing method, wherein the step of inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature, specifically includes:
[0031] Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module to perform element-by-element summation to obtain a first sum feature;
[0032] Inputting the first summed feature into the global context processing subunit of the two-step attention fusion unit and the local context processing subunit of the two-step attention fusion unit for processing at different levels to obtain a first global context processing result and a first local context processing result;
[0033] Sum the first global context processing result and the first local context processing result element by element to obtain a second sum feature, and map the second sum feature through a sigmoid layer to obtain a first mapping feature;
[0034] Multiplying the first mapping feature with the first global feature and the first local feature respectively, and summing the multiplication results element by element to obtain a third sum feature;
[0035] Inputting the third summed feature again into the global context processing subunit and the local context processing subunit for processing at different levels to obtain a second global context processing result and a second local context processing result;
[0036] Sum the second global context processing result and the second local context processing result element by element to obtain a fourth sum feature, and map the fourth sum feature again through a sigmoid layer to obtain a second mapping feature;
[0037] The second mapping feature is multiplied with the first global feature and the first local feature respectively, and the multiplication results are summed element by element to obtain a first output feature.
[0038] Optionally, in the multi-attention based image processing method, the global context processing subunit comprises: a global average pooling layer, a first point-by-point convolution layer, a first ReLU activation function layer, and a second point-by-point convolution layer;
[0039] The local context processing subunit includes: a third point-by-point convolution layer, a second ReLU activation function layer and a fourth point-by-point convolution layer.
[0040] In addition, to achieve the above-mentioned purpose, the present invention further provides an image processing system based on multi-attention, wherein the image processing system based on multi-attention includes:
[0041] A target image acquisition module is used to acquire an image to be processed, and pre-process the image to be processed to obtain a target image;
[0042] A target model building module is used to build an image processing model based on a dual-stage attention fusion mechanism, train and test the image processing model, and obtain a target model;
[0043] The feature extraction and classification module is used to input the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed.
[0044] In addition, to achieve the above-mentioned purpose, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a multi-attention-based image processing program stored in the memory and executable on the processor, and when the multi-attention-based image processing program is executed by the processor, the steps of the multi-attention-based image processing method as described above are implemented.
[0045] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-attention-based image processing program, and when the multi-attention-based image processing program is executed by a processor, the steps of the multi-attention-based image processing method as described above are implemented.
[0046] In the present invention, an image to be processed is obtained, and the image to be processed is preprocessed to obtain a target image; an image processing model based on a two-stage attention fusion mechanism is constructed, and the image processing model is trained and tested to obtain a target model; the target image is input into the target model for feature extraction and classification to obtain a classification result of the image to be processed. The present invention proposes a new module based on dual attention, which uses multi-head self-attention to enhance the global feature extraction capability of a convolutional neural network, and a convolutional block attention module to enhance local feature extraction, and then fuses features of different levels together through two-step attention fusion, and finally uses the architecture of ResNet to construct an image processing model based on a two-stage attention fusion mechanism, thereby improving the accuracy of image processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is a flow chart of a preferred embodiment of the image processing method based on multi-attention of the present invention;
[0048] Figure 2 It is an overall architecture diagram of the target model in the multi-attention based image processing method of the present invention;
[0049] Figure 3 It is a structural diagram of the MFF Block submodule in the multi-attention based image processing method of the present invention;
[0050] Figure 4 It is a structural diagram of the global feature path in the multi-attention based image processing method of the present invention;
[0051] Figure 5It is a structural diagram of the local feature path in the multi-attention based image processing method of the present invention;
[0052] Figure 6 It is a structural diagram of a dual-step attention fusion unit in the multi-attention based image processing method of the present invention;
[0053] Figure 7 is a structural diagram of a preferred embodiment of the image processing system based on multi-attention of the present invention;
[0054] Figure 8 It is a structural diagram of a preferred embodiment of the terminal of the present invention. DETAILED DESCRIPTION
[0055] The present application provides an image processing method, system and terminal based on multi-attention. In order to make the purpose, technical solution and effect of the present application clearer and more specific, the present application is further described in detail with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0056] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as those generally understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with those in the context of the prior art, and will not be interpreted with idealized or overly formal meanings unless specifically defined as here.
[0057] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of the present invention, the descriptions of "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or suggesting their relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In addition, the technical solutions between the various embodiments can be combined with each other, but they must be based on the ability of ordinary technicians in the field to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.
[0058] The image processing method based on multi-attention described in the preferred embodiment of the present invention is as follows: Figure 1 As shown, the multi-attention based image processing method comprises the following steps:
[0059] Step S10: Acquire an image to be processed, and preprocess the image to be processed to obtain a target image.
[0060] Specifically, an image to be processed (such as a medical image) is obtained, and the MRI image is preprocessed. The preprocessing process includes preprocessing techniques such as random cropping, data cleaning, and data normalization, and finally a target image for inputting a model is obtained.
[0061] Step S20: construct an image processing model based on a dual-stage attention fusion mechanism, train and test the image processing model, and obtain a target model.
[0062] In this embodiment, a new dual stage attention fusion mechanism (DSAF) is designed, and an image processing model is constructed based on the dual stage attention fusion mechanism. The image processing model combines convolution operation, multi-head self attention (MHSA) mechanism, convolutional block attention (CBAM) mechanism and two-step attention fusion (DSAF) mechanism, which can enhance the feature representation capability, dynamically adjust the weight distribution of global and local features, and significantly improve the detection capability of complex areas in medical image classification tasks.
[0063] Furthermore, the training and testing of the image processing model also includes:
[0064] Acquiring historical data, the historical data including historical image data and classification results corresponding to the historical image data;
[0065] Using the historical image data as training samples and the classification results corresponding to the historical image data as labels to construct a data set;
[0066] The data set is divided into a training set, a test set and a validation set according to a preset ratio. The training set is used to train the image processing model, the test set is used to evaluate the image processing model in each round of training, and the validation set is used to evaluate the trained image processing model.
[0067] like Figure 2 As shown, further, the target model includes: a first scale feature extraction module, a second scale feature extraction module, a third scale feature extraction module, a fourth scale feature extraction module, a global average pooling layer and a classifier;
[0068] Among them, the first scale feature extraction module includes a convolution layer and an MFF Block submodule, and the second scale feature extraction module, the third scale feature extraction module and the fourth scale feature extraction module all include an average pooling layer and an MFF Block submodule.
[0069] It can be understood that when the image is input to the target model for calculation, it includes four stages (Stage1 to Stage4), and each scale feature extraction module corresponds to a stage. In each stage, this module is repeated many times, similar to the residual module design in ResNet, which can effectively extract local and global information at multiple scales. Specifically, Stage1 is repeated 3 times, the feature map size is 56×56×256, the Stage2 feature map size becomes 28×28×512, repeated 4 times, Stage3 reduces the feature map to 14×14×1024, repeated 6 times, and finally Stage4 reduces it to 7×7×2048, repeated 3 times. Finally, the network generates a vector for classification after dimensionality reduction through the global average pooling layer (AvgPool), and the final result is output through the classifier. Through repeated modular design and combined with the residual connection of ResNet, the robustness of feature extraction and the classification ability of the model are enhanced.
[0070] Step S30: input the target image into the target model for feature extraction and classification to obtain a classification result of the image to be processed.
[0071] Specifically, the target image is input into the convolution layer of the first scale feature extraction module for convolution to obtain a first input feature, the first input feature is input into the global feature path of the MFF Block submodule for feature extraction to obtain a first global feature, and the first input feature is input into the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first local feature.
[0072] It is understandable that if Figure 3As shown, in order to improve the accuracy of the medical image classification model, this embodiment proposes a feature extraction module MFF Block that integrates multiple attention mechanisms. The module combines convolution operations, multi-head self-attention (MHSA) and attention mechanisms to enhance feature representation capabilities. First, the input features are reduced in dimension by 1x1 convolution to reduce computational complexity. Then the features are processed through two parallel paths. In the local feature path, the features are first extracted through 3x3 convolution, and then through CBAM, which includes channel attention and spatial attention mechanisms, respectively emphasizing important channel features and key spatial positions, and finally the number of channels is restored to the input size through 1x1 convolution. In the global feature path, the feature map passes through the MHSA mechanism to capture the dependencies between global features, and restores the original number of channels through 1x1 convolution. Finally, the outputs of the two paths are fused together through the DSAF module to integrate local features, global features, and information extracted by different attention mechanisms. The entire module also contains residual connections, which directly add input features to fused features, thereby enhancing the feature representation capabilities of the model while retaining the original information. This module is designed to improve the performance of the model in various visual tasks.
[0073] The step of inputting the first input feature to the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first global feature specifically includes:
[0074] Input the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module to perform an all2all attention calculation, and linearly transform the first input feature with a query weight matrix, a keyword weight matrix, and a position encoding weight matrix, respectively, to obtain a query vector, a key vector, and a value vector;
[0075] Encode the height and width of the first input feature to obtain a position encoding matrix, multiply the position encoding matrix by the query vector to obtain a first intermediate vector, and multiply the key vector by the query vector to obtain a second intermediate vector;
[0076] The first intermediate vector and the second intermediate vector are summed element by element, the sum result is input into the softmax layer for normalization to obtain a normalized result, and the normalized result is multiplied by the value vector to obtain a first global feature.
[0077] like Figure 4As shown, it is understandable that imaging methods and clinical pathological features in medical images show significant diversity, including significant intra-class variation (for example, the same type of lesion may appear differently in different patients or different parts of the same patient) and inter-class similarity (for example, different types of lesions may look very similar in images). In this case, classification or segmentation based solely on local features may not be sufficient to accurately distinguish different types of lesions. Therefore, it becomes crucial to obtain global semantic information, because global features help the model better understand and distinguish complex features. In order to enhance the model's ability to capture global semantic information, the MHSA mechanism is introduced in the global feature extraction branch. MHSA calculates the relationship between different positions in the feature map, effectively captures global dependencies and enhances the model's representation of complex lesions. This global feature extraction is particularly useful when processing medical images with significant intra-class differences and high inter-class similarity, which can significantly improve the model's discriminative ability and robustness.
[0078] Specifically, the first input feature is input into the head, the head is used to perform all2all attention on the 2D feature map, and the height R of the first input feature is h and width R w Encode to obtain the position encoding matrix r. The output of the self-attention mechanism of a single head can be expressed as:
[0079]
[0080] Among them, O head represents the output of the self-attention mechanism of the head, Softmax() represents the normalization function, X represents the first input feature, and W q , W k and W v They represent the query weight matrix, keyword weight matrix and position encoding weight matrix respectively, and T represents the matrix transpose. q That is the query vector q, XW k That is the key vector k, XW v That is the value vector v.
[0081] Finally, the outputs of all heads are connected and projected again, which can be expressed as:
[0082] MHA(X) = Concat[O1, ...O head , ... N ]W 0 ;
[0083] Among them, MHA(X) represents the value after the output of all heads are connected and projected again, W 0 represents the learned linear transformation, O1 represents the first head, O NIndicates the last header, N indicates the number of headers, and Concat[] indicates the concatenation operation.
[0084] The step of inputting the first input feature to the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first local feature specifically includes:
[0085] Inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module to perform a maximum pooling operation and an average pooling operation to obtain a first maximum pooling result and a first average pooling result;
[0086] Inputting the first maximum pooling result and the first average pooling result into a shared multi-layer perceptron for feature extraction to obtain a first intermediate feature and a second intermediate feature;
[0087] The first intermediate feature and the second intermediate feature are summed element by element, and the summation result is passed through a sigmoid layer to obtain channel attention;
[0088] Performing a maximum pooling operation and an average pooling operation on the channel attention to obtain a second maximum pooling result and a second average pooling result;
[0089] A feature expression of the target local area is extracted from the second maximum pooling result and the second average pooling result, and the feature expression is passed through a sigmoid layer to obtain spatial attention, and the spatial attention is used as the first local feature.
[0090] It should be noted that local features in images are equally important. They can capture the details and local structures of the region and are the key to accurate positioning and classification. In the image analysis process, local feature extraction helps the model focus on the microscopic performance of the target area, such as tissue texture, edge contours, and morphology of local areas, which are often necessary to distinguish different types of images.
[0091] like Figure 5 As shown, in this embodiment, in order to make full use of local features, CBAM is added to the model design. The CBAM module enhances the model's ability to capture local features through two steps: first, the channel attention module assigns different weights to different feature channels, enhances the model's attention to features with diagnostic value, and improves its ability to distinguish different lesion features; next, the spatial attention module captures the relationship between different spatial positions in the image, highlights the feature expression of local areas, and thus helps the model identify and locate lesions in complex backgrounds.
[0092] Furthermore, the first global feature and the first local feature are input into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature.
[0093] Specifically, the first global feature and the first local feature are input into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for element-by-element summation to obtain a first sum feature; the first sum feature is input into the global context processing subunit of the two-step attention fusion unit and the local context processing subunit of the two-step attention fusion unit for processing at different levels to obtain a first global context processing result and a first local context processing result;
[0094] The first global context processing result and the first local context processing result are summed element by element to obtain a second sum feature, and the second sum feature is mapped through a sigmoid layer to obtain a first mapping feature; the first mapping feature is multiplied with the first global feature and the first local feature respectively, and the multiplication results are summed element by element to obtain a third sum feature; the third sum feature is input again into the global context processing subunit and the local context processing subunit for processing at different levels to obtain a second global context processing result and a second local context processing result;
[0095] The second global context processing result and the second local context processing result are summed element by element to obtain a fourth sum feature, and the fourth sum feature is mapped again through a sigmoid layer to obtain a second mapping feature; the second mapping feature is multiplied with the first global feature and the first local feature respectively, and the multiplication results are summed element by element to obtain a first output feature.
[0096] like Figure 6As shown, it can be understood that in the global context and local context processing modules, the features are processed at different levels to obtain global features and local features, respectively. After processing, the obtained features are added again to obtain channel attention across multiple scales. Next, it is necessary to apply the Sigmoid activation function to the calculated channel attention so that its value falls between 0 and 1, and then multiply these channel attentions with the global features and local features respectively by element-by-element multiplication. As the input of the attention module, the initial fusion quality will greatly affect the final fusion weight. Since this is still a feature fusion problem, a novel approach is to use another attention module to fuse the input features, calling this two-stage method dual-step attention fusion. DSAF significantly improves the accuracy of feature fusion and the adaptability of the model by introducing a two-stage attention mechanism.
[0097] DSAF effectively integrates global and local features through multiple fusions, reducing the impact of initial fusion inaccuracies. In addition, DSAF dynamically adjusts fusion weights, enhancing the expressiveness and robustness of the model when processing complex tasks, making it perform well in complex scenarios such as image analysis.
[0098] The global context processing subunit includes: a global average pooling layer, a first point-by-point convolution layer, a first ReLU activation function layer, and a second point-by-point convolution layer; the local context processing subunit includes: a third point-by-point convolution layer, a second ReLU activation function layer, and a fourth point-by-point convolution layer (such as Figure 6 shown).
[0099] Furthermore, the first output feature is input into the average pooling layer of the second scale feature extraction module for average pooling to obtain a second input feature, the second input feature is input into the global feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second global feature, and the first output feature is input into the local feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second local feature.
[0100] The second global feature and the second local feature are input into the two-step attention fusion unit of the MFF Block submodule of the second scale feature extraction module for feature fusion to obtain a second output feature, and the second global feature is continuously input into the next scale feature extraction module for feature extraction and fusion until the target output feature is obtained after passing through the fourth scale feature extraction module.
[0101] The target output feature is input into the global average pooling layer for dimensionality reduction to obtain a vector for classification, and after the vector passes through the classifier, the classification result of the image to be processed is obtained.
[0102] The present invention proposes a dual-branch parallel hierarchical fusion network structure, the main advantages of which are as follows:
[0103] (1) Efficient fusion of global and local features: This paper designs a new dual-stage attention fusion mechanism (DSAF), which can dynamically adjust the weight distribution of global and local features, significantly improving the detection capability of complex lesion areas in medical image classification tasks.
[0104] (2) Combining multi-head self-attention and convolutional block attention modules: The present invention introduces MHSA in global feature extraction and adopts CBAM in local feature extraction, effectively combining the advantages of Transformer and CNN, and enhancing the robustness of the model while maintaining high-precision classification.
[0105] (3) Modular architecture design: The present invention proposes a multi-level feature extraction module MFF Block, which improves the ability to capture multi-level features of images by combining multi-scale feature extraction with residual connection.
[0106] (4) Optimization for medical image classification tasks: The method of the present invention is optimized specifically for the characteristics of high similarity between medical image categories and large differences within categories, so that it can still accurately locate the target (lesion) area under complex backgrounds.
[0107] Furthermore, if Figure 7 As shown, based on the above-mentioned multi-attention-based image processing method, the present invention also provides a multi-attention-based image processing system, wherein the multi-attention-based image processing system includes:
[0108] The target image acquisition module 51 is used to acquire the image to be processed, and pre-process the image to be processed to obtain the target image;
[0109] A target model building module 52 is used to build an image processing model based on a dual-stage attention fusion mechanism, train and test the image processing model, and obtain a target model;
[0110] The feature extraction and classification module 53 is used to input the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed.
[0111] Furthermore, if Figure 8As shown, based on the above-mentioned multi-attention-based image processing method and system, the present invention also provides a terminal accordingly, and the terminal includes a processor 10, a memory 20 and a display 30. Figure 8 Only some components of the terminal are shown, but it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.
[0112] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory of the terminal. In other embodiments, the memory 20 may also be an external storage device of the terminal, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (SecureDigital, SD) card, a flash card (Flash Card), etc. equipped on the terminal. Further, the memory 20 may also include both an internal storage unit of the terminal and an external storage device. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code of the installation terminal. The memory 20 may also be used to temporarily store data that has been output or is to be output. In one embodiment, a multi-attention-based image processing program 40 is stored on the memory 20, and the multi-attention-based image processing program 40 can be executed by the processor 10, thereby realizing the multi-attention-based image processing method in the present application.
[0113] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor or other data processing chip, used to run the program code or process data stored in the memory 20, such as executing the multi-attention-based image processing method.
[0114] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 30 is used to display information on the terminal and to display a visual user interface. The components of the terminal communicate with each other via a system bus.
[0115] In one embodiment, when the processor 10 executes the multi-attention based image processing program 40 in the memory 20, the following steps are implemented:
[0116] Acquire an image to be processed, and preprocess the image to be processed to obtain a target image;
[0117] Constructing an image processing model based on a dual-stage attention fusion mechanism, training and testing the image processing model to obtain a target model;
[0118] The target image is input into the target model for feature extraction and classification to obtain a classification result of the image to be processed.
[0119] The target model includes: a first scale feature extraction module, a second scale feature extraction module, a third scale feature extraction module, a fourth scale feature extraction module, a global average pooling layer and a classifier;
[0120] Among them, the first scale feature extraction module includes a convolution layer and an MFF Block submodule, and the second scale feature extraction module, the third scale feature extraction module and the fourth scale feature extraction module all include an average pooling layer and an MFF Block submodule.
[0121] The MFF Block submodule includes: a global feature path, a local feature path, and a two-step attention fusion unit;
[0122] The step of inputting the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed specifically includes:
[0123] Inputting the target image into the convolution layer of the first scale feature extraction module for convolution to obtain a first input feature, inputting the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first global feature, and inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first local feature;
[0124] Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature;
[0125] Inputting the first output feature to the average pooling layer of the second scale feature extraction module for average pooling to obtain a second input feature, inputting the second input feature to the global feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second global feature, and inputting the first output feature to the local feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second local feature;
[0126] Inputting the second global feature and the second local feature into the two-step attention fusion unit of the MFF Block submodule of the second scale feature extraction module for feature fusion to obtain a second output feature, and continuously inputting the second global feature into the next scale feature extraction module for feature extraction and fusion, until the target output feature is obtained after passing through the fourth scale feature extraction module;
[0127] The target output feature is input into the global average pooling layer for dimensionality reduction to obtain a vector for classification, and after the vector passes through the classifier, the classification result of the image to be processed is obtained.
[0128] The step of inputting the first input feature to the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first global feature specifically includes:
[0129] Input the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module to perform an all2all attention calculation, and linearly transform the first input feature with a query weight matrix, a keyword weight matrix, and a position encoding weight matrix, respectively, to obtain a query vector, a key vector, and a value vector;
[0130] Encode the height and width of the first input feature to obtain a position encoding matrix, multiply the position encoding matrix by the query vector to obtain a first intermediate vector, and multiply the key vector by the query vector to obtain a second intermediate vector;
[0131] The first intermediate vector and the second intermediate vector are summed element by element, the sum result is input into the softmax layer for normalization to obtain a normalized result, and the normalized result is multiplied by the value vector to obtain a first global feature.
[0132] The step of inputting the first input feature to the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first local feature specifically includes:
[0133] Inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module to perform a maximum pooling operation and an average pooling operation to obtain a first maximum pooling result and a first average pooling result;
[0134] Inputting the first maximum pooling result and the first average pooling result into a shared multi-layer perceptron for feature extraction to obtain a first intermediate feature and a second intermediate feature;
[0135] The first intermediate feature and the second intermediate feature are summed element by element, and the summation result is passed through a sigmoid layer to obtain channel attention;
[0136] Performing a maximum pooling operation and an average pooling operation on the channel attention to obtain a second maximum pooling result and a second average pooling result;
[0137] A feature expression of the target local area is extracted from the second maximum pooling result and the second average pooling result, and the feature expression is passed through a sigmoid layer to obtain spatial attention, and the spatial attention is used as the first local feature.
[0138] The step of inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature specifically includes:
[0139] Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module to perform element-by-element summation to obtain a first sum feature;
[0140] Inputting the first summed feature into the global context processing subunit of the two-step attention fusion unit and the local context processing subunit of the two-step attention fusion unit for processing at different levels to obtain a first global context processing result and a first local context processing result;
[0141] Sum the first global context processing result and the first local context processing result element by element to obtain a second sum feature, and map the second sum feature through a sigmoid layer to obtain a first mapping feature;
[0142] Multiplying the first mapping feature with the first global feature and the first local feature respectively, and summing the multiplication results element by element to obtain a third sum feature;
[0143] Inputting the third summed feature again into the global context processing subunit and the local context processing subunit for processing at different levels to obtain a second global context processing result and a second local context processing result;
[0144] Sum the second global context processing result and the second local context processing result element by element to obtain a fourth sum feature, and map the fourth sum feature again through a sigmoid layer to obtain a second mapping feature;
[0145] The second mapping feature is multiplied with the first global feature and the first local feature respectively, and the multiplication results are summed element by element to obtain a first output feature.
[0146] The global context processing subunit includes: a global average pooling layer, a first point-by-point convolution layer, a first ReLU activation function layer, and a second point-by-point convolution layer;
[0147] The local context processing subunit includes: a third point-by-point convolution layer, a second ReLU activation function layer and a fourth point-by-point convolution layer.
[0148] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a multi-attention-based image processing program, and when the multi-attention-based image processing program is executed by a processor, the steps of the multi-attention-based image processing method as described above are implemented.
[0149] In summary, the present invention proposes an image processing method, system and terminal based on multi-attention, the method comprising: obtaining an image to be processed, preprocessing the image to be processed to obtain a target image; constructing an image processing model based on a two-stage attention fusion mechanism, training and testing the image processing model to obtain a target model; inputting the target image into the target model for feature extraction and classification to obtain a classification result of the image to be processed. The present invention proposes a new module based on dual attention, which uses multi-head self-attention to enhance the global feature extraction capability of the convolutional neural network, while the convolutional block attention module enhances local feature extraction, and then fuses features of different levels together through two-step attention fusion, and finally uses the architecture of ResNet to construct an image processing model based on a two-stage attention fusion mechanism, thereby improving the accuracy of image processing.
[0150] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or terminal including the element.
[0151] Of course, those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0152] It should be understood that the application of the present invention is not limited to the above examples. For ordinary technicians in this field, improvements or changes can be made based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to the present invention.
Claims
1. A multi-attention based image processing method, characterized in that: The multi-attention based image processing method comprises: Acquire an image to be processed, and preprocess the image to be processed to obtain a target image; Constructing an image processing model based on a dual-stage attention fusion mechanism, training and testing the image processing model to obtain a target model; The target image is input into the target model for feature extraction and classification to obtain a classification result of the image to be processed.
2. The multi-attention based image processing method according to claim 1, characterized in that: The target model includes: a first scale feature extraction module, a second scale feature extraction module, a third scale feature extraction module, a fourth scale feature extraction module, a global average pooling layer and a classifier; Among them, the first scale feature extraction module includes a convolution layer and an MFF Block submodule, and the second scale feature extraction module, the third scale feature extraction module and the fourth scale feature extraction module all include an average pooling layer and an MFF Block submodule.
3. The multi-attention based image processing method according to claim 2, characterized in that: The MFF Block submodule includes: a global feature path, a local feature path and a two-step attention fusion unit; The step of inputting the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed specifically includes: Inputting the target image into the convolution layer of the first scale feature extraction module for convolution to obtain a first input feature, inputting the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first global feature, and inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first local feature; Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFFBlock submodule of the first scale feature extraction module for feature fusion to obtain a first output feature; Input the first output feature to the average pooling layer of the second scale feature extraction module for average pooling to obtain a second input feature, input the second input feature to the global feature path of the MFFBlock submodule of the second scale feature extraction module for feature extraction to obtain a second global feature, and input the first output feature to the local feature path of the MFF Block submodule of the second scale feature extraction module for feature extraction to obtain a second local feature; Inputting the second global feature and the second local feature into the two-step attention fusion unit of the MFFBlock submodule of the second scale feature extraction module for feature fusion to obtain a second output feature, and continuously inputting the second global feature into the next scale feature extraction module for feature extraction and fusion, until the target output feature is obtained after passing through the fourth scale feature extraction module; The target output feature is input into the global average pooling layer for dimensionality reduction to obtain a vector for classification, and after the vector passes through the classifier, the classification result of the image to be processed is obtained.
4. The multi-attention based image processing method according to claim 3, characterized in that: Inputting the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain a first global feature specifically includes: Input the first input feature into the global feature path of the MFF Block submodule of the first scale feature extraction module to perform an all2all attention calculation, and linearly transform the first input feature with a query weight matrix, a keyword weight matrix, and a position encoding weight matrix, respectively, to obtain a query vector, a key vector, and a value vector; Encode the height and width of the first input feature to obtain a position encoding matrix, multiply the position encoding matrix by the query vector to obtain a first intermediate vector, and multiply the key vector by the query vector to obtain a second intermediate vector; The first intermediate vector and the second intermediate vector are summed element by element, the sum result is input into the softmax layer for normalization to obtain a normalized result, and the normalized result is multiplied by the value vector to obtain a first global feature.
5. The multi-attention based image processing method according to claim 3, characterized in that: Inputting the first input feature to the local feature path of the MFF Block submodule of the first scale feature extraction module for feature extraction to obtain the first local feature specifically includes: Inputting the first input feature into the local feature path of the MFF Block submodule of the first scale feature extraction module to perform a maximum pooling operation and an average pooling operation to obtain a first maximum pooling result and a first average pooling result; Inputting the first maximum pooling result and the first average pooling result into a shared multi-layer perceptron for feature extraction to obtain a first intermediate feature and a second intermediate feature; The first intermediate feature and the second intermediate feature are summed element by element, and the summation result is passed through a sigmoid layer to obtain channel attention; Performing a maximum pooling operation and an average pooling operation on the channel attention to obtain a second maximum pooling result and a second average pooling result; A feature expression of the target local area is extracted from the second maximum pooling result and the second average pooling result, and the feature expression is passed through a sigmoid layer to obtain spatial attention, and the spatial attention is used as the first local feature.
6. The multi-attention based image processing method according to claim 3, characterized in that: The step of inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFF Block submodule of the first scale feature extraction module for feature fusion to obtain a first output feature specifically includes: Inputting the first global feature and the first local feature into the two-step attention fusion unit of the MFFBlock submodule of the first scale feature extraction module to perform element-by-element summation to obtain a first sum feature; Inputting the first summed feature into the global context processing subunit of the two-step attention fusion unit and the local context processing subunit of the two-step attention fusion unit for processing at different levels to obtain a first global context processing result and a first local context processing result; Sum the first global context processing result and the first local context processing result element by element to obtain a second sum feature, and map the second sum feature through a sigmoid layer to obtain a first mapping feature; Multiplying the first mapping feature with the first global feature and the first local feature respectively, and summing the multiplication results element by element to obtain a third sum feature; Inputting the third summed feature again into the global context processing subunit and the local context processing subunit for processing at different levels to obtain a second global context processing result and a second local context processing result; Sum the second global context processing result and the second local context processing result element by element to obtain a fourth sum feature, and map the fourth sum feature again through a sigmoid layer to obtain a second mapping feature; The second mapping feature is multiplied with the first global feature and the first local feature respectively, and the multiplication results are summed element by element to obtain a first output feature.
7. The multi-attention based image processing method according to claim 6, characterized in that: The global context processing subunit includes: a global average pooling layer, a first point-by-point convolution layer, a first ReLU activation function layer, and a second point-by-point convolution layer; The local context processing subunit includes: a third point-by-point convolution layer, a second ReLU activation function layer and a fourth point-by-point convolution layer.
8. A multi-attention based image processing system, characterized in that: The multi-attention based image processing system comprises: A target image acquisition module is used to acquire an image to be processed, and pre-process the image to be processed to obtain a target image; A target model building module is used to build an image processing model based on a dual-stage attention fusion mechanism, train and test the image processing model, and obtain a target model; The feature extraction and classification module is used to input the target image into the target model for feature extraction and classification to obtain the classification result of the image to be processed.
9. A terminal, characterized in that: The terminal includes: a memory, a processor, and a multi-attention-based image processing program stored in the memory and executable on the processor. When the multi-attention-based image processing program is executed by the processor, the steps of the multi-attention-based image processing method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a multi-attention-based image processing program, and when the multi-attention-based image processing program is executed by a processor, the steps of the multi-attention-based image processing method as described in any one of claims 1-7 are implemented.