Image processing method and device based on attention, electronic equipment and storage medium
Through the attention-based image processing method, self-attention feature extraction and feature slice fusion technology, the problem of low segmentation accuracy of lesion areas in MRI images is solved, and high-precision segmentation and recognition of abnormal areas in MRI images is achieved.
Patent Information
- Application Number
- CN202311539630.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-18
- Publication Date
- 2025-07-08
AI Technical Summary
Existing deep learning algorithms fail to fully pay attention to the most valuable pixel points in MRI images, resulting in low accuracy of lesion area segmentation and easy loss of important information during feature fusion.
The attention-based image processing method is adopted to improve the recognition ability and segmentation accuracy of the lesion area through step-by-step multi-layer feature extraction, self-attention feature extraction and feature enhancement processing, feature slice fusion and feature guidance processing.
Through self-attention feature mechanism and feature slice fusion processing, semantic information is effectively preserved, and the segmentation accuracy and recognition ability of abnormal areas in MRI images are improved.
Smart Images

Figure CN120279261A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and particularly to an attention-based image processing method, apparatus, electronic device, and storage medium. Background Art
[0002] Currently, the automatic segmentation of three-dimensional medical images plays an important guiding role in clinical diagnosis and treatment. The segmentation and detection of abnormal regions in magnetic resonance images (MRI images) are beneficial to assisting the establishment of clinical treatment decisions, and at the same time can assist in improving the accuracy of surgical treatment and the postoperative observation ability. In the traditional fields of image segmentation and abnormal region recognition, if one wants to segment an abnormal lesion region (such as a uterine fibroids lesion region) from, for example, an MRI image, prior knowledge is often required, such as the shape of uterine fibroids and the characteristics of the uterine fibroids region. However, uterine fibroids have complex and variable shapes and can appear at any position in the uterus, which limits the application of traditional segmentation methods in clinical practice. In recent years, data-driven and automatically feature-extracting deep learning algorithms have achieved great success in the field of computer vision. In particular, convolutional neural networks have achieved better segmentation effects than traditional methods in natural light image segmentation tasks. However, current deep learning algorithms fail to fully focus on the most valuable pixel points in MRI images and are prone to losing important information during the feature fusion process, resulting in the need to improve the segmentation accuracy of lesion regions. Summary of the Invention
[0003] Embodiments of the present disclosure provide an attention-based image processing method, apparatus, electronic device, and storage medium, and the present disclosure can improve the recognition ability of lesion regions and the segmentation accuracy of lesions.
[0004] According to a first aspect of the present disclosure, there is provided an attention-based image processing method, the method including:
[0005] Performing hierarchical multi-level feature extraction on an input image to respectively obtain a plurality of corresponding basic feature maps, the basic feature maps including low-level features and high-level features;
[0006] Performing self-attention feature extraction and feature enhancement processing on the basic feature maps in sequence to respectively obtain enhanced feature maps corresponding to the basic feature maps;
[0007] Performing feature slice fusion processing on the enhanced feature map of the high-level features to obtain high-level attention coefficients and a high-level segmentation feature map;
[0008] Performing feature guidance processing on the enhanced feature map of the low-level features by using the high-level attention coefficients to obtain a guidance fusion feature map;
[0009] Perform anomaly detection using the high-level segmentation feature map and the guidance fusion feature map to obtain an anomaly detection result.
[0010] In some possible implementation manners, the sequentially performing self-attention feature extraction and feature enhancement processing on the base feature map to respectively obtain an enhanced feature map corresponding to the base feature map includes:
[0011] Perform channel attention feature extraction and location attention information mining on the base feature map respectively to correspondingly obtain a channel attention feature and location attention information;
[0012] Perform element-wise addition processing on the channel attention feature and the location attention information to obtain a primary feature map;
[0013] Perform a feature enhancement operation on the primary feature map to obtain an enhanced feature map corresponding to the base feature map.
[0014] In some possible implementation manners, the deep features include front-layer features, middle-layer features, and back-layer features;
[0015] The performing feature slice fusion processing on the enhanced feature map of the deep features to obtain a high-level attention coefficient and a high-level segmentation feature map includes:
[0016] Perform an upsampling operation on the enhanced feature map of the middle-layer features and perform an element-wise addition operation with the enhanced feature map of the front-layer features to obtain a high-level fusion feature map;
[0017] Perform slice fusion processing on the high-level fusion feature map using the enhanced feature map of the back-layer features to obtain the high-level attention coefficient and the high-level segmentation feature map.
[0018] In some possible implementation manners, the performing slice fusion processing on the high-level fusion feature map using the enhanced feature map of the back-layer features of the deep features to obtain the high-level attention coefficient and the high-level segmentation feature map includes:
[0019] Perform slice processing on the high-level fusion feature map to obtain a plurality of slice layers;
[0020] Insert the enhanced feature map of the back-layer features between two adjacent slice layers and perform channel merging processing to obtain a first merged feature;
[0021] Perform base convolution processing on the first merged feature and add the result of the base convolution processing to the high-level fusion feature to obtain a second merged feature;
[0022] Perform the base convolution processing based on the second merged feature to obtain a third merged feature;
[0023] The high-level attention coefficient and the high-level segmentation feature map are obtained by using the second merging feature and the third merging feature.
[0024] In some possible implementation manners, the performing feature guidance processing on the enhanced feature map of the low-level features by using the high-level attention coefficient to obtain a guidance fusion feature map includes:
[0025] Performing an element-wise multiplication operation on the high-level attention coefficient and the enhanced feature map of the low-level features respectively to obtain a refined feature map corresponding to the low-level features;
[0026] Performing an element-wise addition operation on the enhanced feature map of the low-level features and the refined feature map respectively to obtain an aggregated feature map;
[0027] The guidance fusion feature map is obtained by performing an addition process on the aggregated features corresponding to the low-level features respectively.
[0028] In some possible implementation manners, the performing anomaly detection by using the high-level segmentation feature map and the guidance fusion feature map to obtain an anomaly detection result includes:
[0029] Performing the feature slice fusion processing on the high-level segmentation feature map and the guidance fusion feature to obtain an image prediction feature;
[0030] Performing anomaly detection on the image feature to obtain the anomaly detection result.
[0031] In some possible implementation manners, the performing anomaly detection on the image feature to obtain the anomaly detection result includes:
[0032] Using a decision tree model, taking the image feature as the root node, and performing anomaly region detection;
[0033] Based on the detected anomaly region, performing anomaly type detection;
[0034] Generating an anomaly detection result based on the detected anomaly region and anomaly type.
[0035] According to a second aspect of the present disclosure, there is provided an attention-based image processing apparatus, the apparatus including:
[0036] A basic feature map extraction module, configured to perform hierarchical multi-level feature extraction on an input image to respectively obtain a plurality of corresponding basic feature maps, the basic feature maps including low-level features and deep features;
[0037] A self-attention feature extraction module, configured to perform self-attention feature extraction and feature enhancement processing on the basic feature map in sequence to respectively obtain an enhanced feature map corresponding to the basic feature map;
[0038] A feature slice fusion module for performing feature slice fusion processing on the enhanced feature map of the deep feature to obtain a high-level attention coefficient and a high-level segmentation feature map;
[0039] A guidance fusion module for performing feature guidance processing on the enhanced feature map of the low-level feature by using the high-level attention coefficient to obtain a guidance fusion feature map;
[0040] An anomaly detection module for performing anomaly detection by using the high-level segmentation feature map and the guidance fusion feature map to obtain an anomaly detection result.
[0041] According to a third aspect of the present disclosure, an electronic device is provided, which includes:
[0042] A processor;
[0043] A memory for storing instructions executable by the processor;
[0044] Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of the first aspect.
[0045] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method according to any one of the first aspect is implemented.
[0046] The present disclosure proposes a self-attention feature mechanism, which can realize in-depth mining of important information and efficiently extract basic feature information generated by a backbone network. In addition, the present disclosure proposes a feature slice fusion processing, by effectively fusing feature slices with deep features, feature information from different layers is obtained, semantic information is maximally retained and noise is suppressed. The method proposed by the present disclosure can improve the segmentation accuracy of the attention area or the anomaly area in the image.
[0047] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure.
[0048] According to the following detailed description of exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. Description of the Drawings
[0049] Figure 1 is a flowchart of an attention-based image processing method according to an embodiment of the present disclosure;
[0050] Figure 2 is a schematic network structure diagram of an attention-based image processing network according to an embodiment of the present disclosure;
[0051] Figure 3 is a flowchart of self-attention feature extraction and feature enhancement processing according to an embodiment of the present disclosure;
[0052] Figure 4 is a schematic diagram of the network structure of a self-attention feature extraction module according to an embodiment of the present disclosure;
[0053] Figure 5 is a schematic diagram of the network structure of a feature enhancement module according to an embodiment of the present disclosure;
[0054] Figure 6 is a schematic diagram of the network structure of a feature slice fusion module according to an embodiment of the present disclosure;
[0055] Figure 7 is a decision tree block diagram of a myoma discrimination and classification module of an example of the present disclosure;
[0056] Figure 8 is a principle block diagram of an attention-based image processing device according to an embodiment of the present disclosure;
[0057] Figure 9 is a block diagram of an electronic device 800 according to an embodiment of the present disclosure;
[0058] Figure 10 is a block diagram of another electronic device 1900 according to an embodiment of the present disclosure. Detailed implementation manners
[0059] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0060] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein does not have to be construed as superior to or better than other embodiments.
[0061] The term "and / or" herein merely describes an association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.
[0062] In addition, to better illustrate the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without certain specific details. In some instances, methods, means, elements, and circuits well-known to those skilled in the art are not described in detail to highlight the gist of the present disclosure.
[0063] The execution subject of the attention-based image processing method according to an embodiment of the present disclosure can be any electronic device. For example, it can be executed by a terminal device, a server, or other processing devices. Among them, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementation manners, the method can be implemented by a processor invoking computer-readable instructions stored in a memory.
[0064] Figure 1 is a flowchart of an attention-based image processing method according to an embodiment of the present disclosure. As shown in FIG. 2 is a schematic diagram of the network structure of an attention-based image processing network according to an embodiment of the present disclosure. As Figure 1 shown, the attention-based image processing method can include:
[0065] S10: Perform hierarchical multi-layer feature extraction on the input image to respectively obtain a plurality of corresponding basic feature maps, where the basic feature maps include low-level features and high-level features;
[0066] In some possible implementation manners, the input image can be any medical image and can be constructed as a three-dimensional medical image; for example, it can include MRI images, computed tomography images (CT images), and the present disclosure does not make specific limitations thereto. In addition, the input image in the embodiments of the present disclosure can be a scanned image of any part, such as the brain, lungs, abdomen, etc. The embodiments of the present disclosure are described by taking a uterine MRI image as an example, and the abnormal area can be a uterine fibroids area, but it is not a specific limitation of the present disclosure.
[0067] In the embodiments of the present disclosure, an MRI image can be obtained using a nuclear magnetic resonance device, or an input image can be transmitted through other electronic devices, or a corresponding input image can be requested from a server. After obtaining the input image, feature extraction processing can be performed on the MRI image to obtain basic features. Among them, the feature extraction method in the embodiments of the present disclosure can include feature extraction of different depths, that is, by setting multi-layer feature extraction, feature information of different depths and levels can be obtained. In one example, the input image can be sequentially subjected to five-layer depth feature extraction using a VGG-16 or Res2Net-50 backbone network to obtain five (five) basic feature maps. In the embodiments of the present disclosure, the five basic feature maps are respectively denoted by the symbols F1 M , F2 M , F3 M , F4 M and F5 M .
[0068] In the embodiments of the present disclosure, the obtained basic feature maps are feature maps of different depths. In the embodiments of the present disclosure, the feature maps of different depths can be divided into deep features and low-level features. For example, the low-level features are the first two basic feature maps, and the deep features are the last three basic feature maps.
[0069] S20: Sequentially perform self-attention feature extraction and feature enhancement processing on the basic feature maps to respectively obtain enhanced feature maps of the basic feature maps;
[0070] In some possible implementation manners, self-attention feature extraction and feature enhancement processing can be performed on the obtained five basic feature maps. Among them, after processing the self-attention feature maps, the primary features corresponding to the basic features can be respectively obtained. In the embodiments of the present disclosure, the primary features corresponding to each layer of basic features can be denoted as F1 SAM , F2 SAM , F3 SAM , F4 SAM and F5 SAM . After performing feature enhancement processing on each primary feature, the corresponding enhanced feature maps can be obtained, which are respectively denoted as F1 FEM , F2 FEM , F3 FEM , F4 FEM and F5 FEM .
[0071] S30: Perform feature slice fusion processing on the enhanced feature maps of the deep features to obtain high-level attention coefficients and high-level segmentation feature maps;
[0072] In some possible implementation manners, feature slice fusion processing can be performed on the enhanced feature maps corresponding to the deep features in the basic features to achieve full fusion of the deep feature information and lay a foundation for the guidance and optimization of the low-level features.
[0073] S40: Perform feature guidance processing on the enhanced feature map of the low-level features using the high-level attention coefficient to obtain a guided fusion feature map;
[0074] In some possible implementation manners, the Sigmoid function can be first used to activate the high-level attention coefficient, and perform element-wise multiplication and addition with the enhanced feature maps corresponding to the low-level features respectively to obtain the aggregated feature maps corresponding to the low-level features. Then, perform an element-wise addition operation on the aggregated feature maps corresponding to the low-level features respectively to obtain a low-level aggregation map, that is, a guided fusion feature map.
[0075] S50: Perform anomaly detection using the high-level segmentation feature map and the guided fusion feature map to obtain an anomaly detection result.
[0076] In some possible implementation manners, a feature slice fusion operation can be performed using the guided fusion feature map and the high-level segmentation feature map to obtain a segmentation prediction map, so as to obtain an anomaly region segmentation. Then, a classification operation can be performed on the segmentation prediction map using a classification model to obtain a judgment of the anomaly degree.
[0077] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. First, the ways for the embodiments of the present disclosure to obtain an input image may include at least one of the following ways:
[0078] A) Directly collect the input image using a medical imaging acquisition device;
[0079] B) Transmit and receive the input image through an electronic device;
[0080] The embodiments of the present disclosure can receive the input image transmitted by other electronic devices through a communication method, and the communication method may include wired communication and / or wireless communication;
[0081] C) Read the input image stored in the database;
[0082] The embodiments of the present disclosure can read the three-month personnel images stored locally or the input images stored in the server according to the received data reading instruction, and the present disclosure does not make specific limitations on this.
[0083] After obtaining the input MRI image, segmentation and anomaly detection of the abnormal lesion area can be performed on the input image. Among them, first, a backbone network can be used to perform multi-level feature extraction on the input image, where the backbone network can be optionally VGG-16 or Res2Net-50, to obtain the first basic feature F1 M , the second basic feature F2 M , the third basic feature
[0084] F3 M, the fourth basic feature F4 M , the fifth basic feature F5 M , providing a basis for the enhancement and fusion of subsequent features.
[0085] After obtaining the basic features at each level, self-attention feature extraction and feature enhancement processing can be performed on the basic features. Figure 3 is a flowchart of self-attention feature extraction and feature enhancement processing according to an embodiment of the present disclosure. As Figure 3 shown, the self-attention feature extraction and feature enhancement processing include:
[0086] S21: Perform channel attention feature extraction and location attention information mining on the basic feature map respectively, and correspondingly obtain channel attention features and location attention information;
[0087] S22: Perform element addition processing on the channel attention features and the location attention information to obtain a primary feature map;
[0088] S23: Perform a feature enhancement operation on the primary feature map to obtain an enhanced feature map corresponding to the basic feature map.
[0089] Figure 4 is a schematic diagram of the network structure of the self-attention feature extraction module according to an embodiment of the present disclosure. Among them, the self-attention feature extraction module SAM is composed of a channel attention feature extraction module PAM and a location attention information mining module CAM. In one embodiment, the five-layer basic features F1 M , F2 M , F3 M , F4 M , F5 M are respectively used as input features and input into the self-attention feature extraction module for attention information extraction. The following takes the fifth-layer basic feature F5 M as an example to illustrate the extraction process. The processing processes of the basic features of other layers are the same and will not be repeated. First, perform channel attention feature extraction on the basic feature F5 M . Perform an adaptive max pooling (AMP) operation with a dimension of 1 on the basic feature F5 M to obtain global important feature information, and then sequentially pass through 4 convolutional blocks with a convolution kernel of 1. The first convolutional block reduces the number of channels to 1 / 4 of the original, the second convolutional block reduces the number of channels to 1 / 4 of the output features of the previous convolutional block again, the third convolutional block expands the number of channels by 4 times, and the fourth convolutional block expands the number of channels by 4 times again, so as to restore the number of channels to the channel information of F5 M . The output of the fourth convolutional block is combined with the basic feature F5 MThe result of multiplying the elements serves as the output of the channel attention feature extraction module, obtaining the channel attention feature F5 corresponding to the base feature PAM . Additionally, the methods for obtaining the channel attention features of the first, second, third, and fourth layers are the same as those of the fifth layer. The computational model for channel attention feature extraction can be expressed as:
[0090]
[0091] where F i PAM represents the obtained channel attention feature, i = 1, 2, 3, 4, 5, and F i M represents the i-th layer base feature, CBR represents the base convolution operation (batch normalization and Relu activation after convolution), AMP represents adaptive max pooling; C×4 means the number of channels is magnified by 4 times. C / 4 means the number of channels is reduced to 1 / 4 of the original; represents element-wise multiplication.
[0092] While extracting the channel attention feature corresponding to the base feature, it is also possible to extract the localization attention information from the base feature F5 M . Specifically, the localization attention information extraction module includes a total of four branches, such as Figure 4 the four branches of the CAM module from top to bottom in. In the first branch, the present disclosure example first converts the base feature into a one-dimensional feature, and the dimension can be expressed as (1, 1, n), where n represents the number of features in the third dimension (the number of feature points in length and width), the first 1 is the number of features in the first dimension, which can represent the batch size, and the second 1 is the number of features in the second dimension. The present disclosure embodiment can use the torch.view(-1) function to change the shape of F5 M to make it a one-dimensional feature matrix. Then, the one-dimensional feature matrix can be transposed, swapping the positions of the second and third dimensions, such as using the.permute(0, 2, 1) function to achieve the matrix transpose, denoted as . Similarly, in the second branch, only the base feature is converted into a one-dimensional feature, such as using the.view(-1) function to change the shape of F5 M to make it a one-dimensional feature matrix, denoted as The third branch is the same as the second branch, only using the.view(-1) function to change the shape of F5 to make it a one-dimensional feature matrix, denoted as The fourth branch is the residual connection branch, which performs matrix multiplication on the transposed matrix obtained from the first branch and the one-dimensional feature matrix obtained from the second branch, and then performs softmax normalization operation to obtain the localization attention coefficient, denoted as Subsequently, it is combined with the one-dimensional feature matrix of the third branch Perform matrix multiplication to obtain the three-branch fusion coefficient, convert this three-branch fusion coefficient into a new one-dimensional feature matrix, such as transposing using.view(-1), and then perform element-wise multiplication of the new one-dimensional feature matrix with the self-learning parameter α. Finally, perform element-wise addition on the multiplication result and the basic feature F5 M Perform element-wise addition to obtain the final location attention information. The computational model for extracting location attention information can be expressed as:
[0093]
[0094]
[0095]
[0096]
[0097]
[0098] Among them, represents the first branch after matrix transposition, represents the second branch after matrix transposition, represents the third branch after matrix transposition, represents the location attention coefficient, F i CAM represents the output of the location attention information module, represents the element-wise multiplication operation, view represents changing the structure of the tensor, permute represents the matrix transposition operation, and α represents the self-learning parameter, which is obtained through network training.
[0099] Based on the above configuration, the channel attention features output by the channel attention feature extraction module can be added to the location attention information output by the location attention information extraction module to obtain the output of the self-attention feature extraction module, that is, the primary feature map, F1 SAM 、F2 SAM 、F3 SAM 、F4 SAM and F5 SAM . Subsequently, feature enhancement operations can be performed on the five-layer primary feature extraction maps to correspondingly obtain five-layer enhanced feature maps.
[0100] Figure 5 is a schematic diagram of the network structure of the feature enhancement module according to an embodiment of the present disclosure. Through the feature enhancement module, feature enhancement operations can be performed on the feature extraction map to obtain the enhanced feature map corresponding to the basic feature map.
[0101] Similarly, below, taking the fifth-layer primary feature map F5 as an exampleSAM Take it as an example to illustrate the extraction process. First, perform the average operation (Mean) and the maximum operation (Max) on the primary feature map F5 respectively SAM Then, concatenate the results of the two in channels. The purpose of this operation is to obtain significant features in the image while maximizing the retention of valid information. Perform the basic convolution operation and the Sigmoid normalization operation on the result after channel concatenation in sequence to obtain the enhanced feature coefficients, and multiply them with the primary feature map F5 SAM to obtain the enhanced feature map F5 FEM . The computational model for feature enhancement can be expressed as:
[0102]
[0103] where represents element-wise multiplication, Sigmoid represents the activation function, CBR represents the basic convolution operation (batch normalization and Relu activation after convolution), Cat represents the channel concatenation operation, Mean represents the average operation; Max represents the maximum operation, and F i FEM represents the output of the feature enhancement module.
[0104] Based on the above configuration, five enhanced feature maps can be obtained corresponding to the five primary feature maps, and then the subsequent feature slice fusion operation can be performed to obtain the high-level attention coefficients and the high-level feature segmentation map.
[0105] In the embodiments of the present disclosure, the deep features include the front-layer features, the middle-layer features, and the back-layer features; performing feature slice fusion processing on the enhanced feature maps of the deep features to obtain the high-level attention coefficients and the high-level segmentation feature maps includes: performing an upsampling operation on the enhanced feature map of the middle-layer features, and performing an element-wise addition operation with the enhanced feature map of the front-layer features to obtain the high-level fusion feature map; using the enhanced feature map of the back-layer features to perform slice fusion processing on the high-level fusion feature map to obtain the high-level attention coefficients and the high-level segmentation feature maps.
[0106] In some possible embodiments, when the basic feature has five layers, the front-layer feature in the deep feature may be the third-layer basic feature, the middle-layer feature may be the fourth-layer basic feature, and the back-layer feature may be the fifth-layer feature. Before slice fusion, first perform the fusion between the enhanced feature map corresponding to the front-layer feature and the enhanced feature map corresponding to the middle-layer feature. Since the dimensions are different, it is necessary to perform upsampling on the enhanced feature map of the middle-layer feature to make the dimensions the same. For example, the nn.Upsample() function can be used to perform bilinear interpolation upsampling on the enhanced feature map of the fourth layer to enlarge the feature map by 2 times. After the upsampling process is completed, add the upsampling result to the enhanced feature map of the front-layer feature to obtain a high-level fusion feature map, that is, perform an element-wise addition operation with the enhanced feature map of the third layer to obtain a high-level fusion feature map. Then, the enhanced feature map of the back-layer feature can be used to perform slice fusion processing on the obtained high-level fusion feature map to obtain the high-level attention coefficient and the high-level segmentation feature map.
[0107] Among them, the use of the enhanced feature map of the back-layer feature of the deep feature to perform slice fusion processing on the high-level fusion feature map includes: performing slice processing on the high-level fusion feature map to obtain multiple slice layers; inserting the enhanced feature map of the back-layer feature between two adjacent slice layers, and performing channel merging processing to obtain a first merged feature; performing basic convolution processing on the first merged feature, and adding the result of the basic convolution processing to the high-level fusion feature to obtain a second merged feature; performing the basic convolution processing based on the second merged feature to obtain a third merged feature; using the second merged feature and the third merged feature to obtain the high-level attention coefficient and the high-level segmentation feature map.
[0108] Figure 6 It is a schematic diagram of the network structure of the feature slice fusion module according to an embodiment of the present disclosure. Through the feature slice fusion module, slice fusion processing can be performed. Specifically, perform slice processing on the high-level fusion feature map to obtain slice layers with a preset number of layers, such as 64 can be set. In practical applications, the torch.chunk() function can be used to complete the slice processing of the high-level fusion feature map to obtain 64 slice layers. Then, the enhanced feature map of the back-layer feature (the fifth layer) can be inserted between two adjacent slice layers. For example, the enhanced feature map of the fifth layer can be inserted after each slice feature. Specifically, the torch.cat() function can be used to perform periodic interpolation operations, and then the first channel merging processing is used to splice the channels of the interpolation result to obtain a first merged feature. Then perform basic convolution (CBR) processing on the first merged feature, and determine the element-wise addition result of the high-level fusion feature map and the result of the basic convolution processing as the second merged feature. After performing CBR processing on the second merged feature, a third merged feature can be obtained.
[0109] In some possible embodiments, the second merged feature may be determined as a high-level segmentation feature map, and the third merged feature may be determined as a high-level attention coefficient. In other embodiments, in order to further extract the feature information most relevant to the target region from the high-level features, multiple slice fusion processes may also be performed, and the number of slices per layer in each slice fusion process is half of that in the previous slice fusion process. For example, the subsequent slice parameters can be set to 32, 16, 8, 4, 2, 1 respectively. After the slice fusion process with a slice parameter of 1, the process ends, and the final high-level segmentation feature map and high-level attention coefficient are obtained.
[0110] Specifically, in the embodiments of the present disclosure, the third merged feature obtained from the first slice fusion process with 64 slices can be defined as λ 64 , and the second merged feature is defined as In the subsequent slice fusion process, the second merged feature can be sliced again using the torch.chunk() function, and periodic interpolation is performed with the third merged feature defined as λ 64 . After multiple slice fusion processes, the high-level attention coefficient λ h and the high-level myoma segmentation map F h fuse are finally obtained. The calculation model for high-level feature slice fusion can be expressed as:
[0111]
[0112]
[0113]
[0114]
[0115]
[0116]
[0117]
[0118]
[0119]
[0120] Among them, FSF represents the feature slice fusion operation, Upsample represents the upsampling operation, and Cat represents the channel concatenation operation.
[0121] After obtaining the high-level attention coefficient, the low-level features can be guided and optimized using the high-level attention coefficient. In the embodiments of the present disclosure, feature guidance processing is performed on the enhanced feature map of the low-level features using the high-level attention coefficient to obtain a guided fusion feature map, including: performing an element-wise multiplication operation between the high-level attention coefficient and the enhanced feature map of the low-level features respectively to obtain a refined feature map corresponding to the low-level features; performing an element-wise addition operation on the enhanced feature map and the refined feature map of the low-level features respectively to obtain an aggregated feature map; and obtaining the guided fusion feature map by performing an addition process on the aggregated feature maps corresponding to the low-level features respectively.
[0122] As Figure 2 shown, in some possible embodiments, the Sigmoid function can be used to perform a normalization operation on the high-level attention coefficient λ h and then perform an element-wise multiplication with the enhanced feature map F1 FEM corresponding to the first-layer basic features and the enhanced feature map F2 FEM corresponding to the second-layer basic features respectively to obtain a first refined feature map and a second refined feature map. Add the two enhanced features to the corresponding refined features respectively to obtain two sets of aggregated feature maps, and finally perform an element-wise addition process on the aggregated feature maps to obtain the guided fusion feature map. In the embodiments of the present disclosure, the two aggregated feature maps can be represented as F1 agg and F2 agg .
[0123] After obtaining the guided fusion feature map, anomaly detection can be performed using the guided fusion feature map and the high-level segmentation feature map. In the embodiments of the present disclosure, anomaly detection is performed using the high-level segmentation feature map and the guided fusion feature map to obtain an anomaly detection result, including: performing the feature slice fusion process using the high-level segmentation feature map and the guided fusion feature to obtain an image prediction feature; and performing anomaly detection on the image feature to obtain the anomaly detection result. The anomaly detection result includes the segmentation of the anomaly region and / or the classification of the anomaly type, the detection of the anomaly degree, etc.
[0124] In the embodiments of the present disclosure, the slice fusion process can be performed using the high-level segmentation feature map and the guided fusion feature map, so as to fully perform feature fusion and key information localization.
[0125] Among them, slicing processing can be performed on the guidance fusion feature to obtain a plurality of slice layers; then, the high-level segmentation feature map is inserted between two adjacent slice layers, and channel merging processing is performed to obtain a fourth merged feature; basic convolution processing is performed on the fourth merged feature, and the result of the basic convolution processing is added to the guidance fusion feature to obtain a fifth merged feature; based on the fifth merged feature, the basic convolution processing is performed to obtain a sixth merged feature; a new high-level attention coefficient and a new high-level segmentation feature map are obtained by using the fifth merged feature and the sixth merged feature.
[0126] Then, the new high-level attention coefficient and the new high-level segmentation feature map are subjected to slice fusion feature processing with different numbers of slice layers until the number of slice layers is 1. The number of slice layers in the latter time is half of that in the previous time, and the new high-level segmentation feature obtained by the last slice fusion processing is determined as the image prediction feature for performing the segmentation of the abnormal region. Then, based on the segmentation result, the classification model can be further used to determine the abnormal classification information. The abnormal classification information may include the type of lesion, the severity of the lesion, etc., and the present disclosure does not make specific limitations thereto.
[0127] Taking uterine fibroids as an example for illustration. Figure 7 It is the decision tree block diagram of the fibroid discrimination and classification module of the present disclosure example. In some possible implementation manners, the image prediction feature is input into the abnormal discrimination and classification module to obtain abnormal classification information, including: using the decision tree model, taking the image prediction feature as the root node, performing classification and discrimination of abnormal features, such as determining the growth site of the fibroid and the relationship between the fibroid and the uterine wall; taking "the growth site of the fibroid" and "the relationship with the uterine wall" as internal nodes for further classification. In the internal node of "the growth site of the fibroid", "corpus uteri fibroid" and "cervical fibroid" are used as leaf nodes for classification; and in the internal node of "the relationship with the uterine wall", "intramural fibroid", "submucous fibroid", and "subserous fibroid" are used as leaf nodes for classification. Based on the above, the detection and discrimination of different abnormal types can be completed. In other implementation manners, the severity, such as early stage, middle stage, and late stage, etc., can also be determined, and the present disclosure does not make specific limitations thereto.
[0128] In addition, the embodiments of the present disclosure also provide an attention-based image processing device, an electronic device, and a storage medium. Among them Figure 8 It is the principle block diagram of the attention-based image processing device according to the embodiment of the present disclosure. Among them, the attention-based image processing device includes:
[0129] A basic feature map extraction module 100, configured to perform hierarchical multi-level feature extraction on the input image to respectively obtain corresponding multiple basic feature maps, where the basic feature maps include low-level features and high-level features;
[0130] The self-attention feature extraction module 200 is used to sequentially perform self-attention feature extraction and feature enhancement processing on the basic feature map to obtain the enhanced feature map corresponding to the basic feature map;
[0131] The feature slice fusion module 300 is used to perform feature slice fusion processing on the enhanced feature map of the deep feature to obtain the high-level attention coefficient and the high-level segmentation feature map;
[0132] The guidance fusion module 400 is used to perform feature guidance processing on the enhanced feature map of the low-level feature by using the high-level attention coefficient to obtain the guidance fusion feature map;
[0133] The anomaly detection module 500 is used to perform anomaly detection by using the high-level segmentation feature map and the guidance fusion feature map to obtain the anomaly detection result.
[0134] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0135] In summary, compared with the prior art, the beneficial effects of the embodiments of the present disclosure include the following aspects:
[0136] 1. The embodiments of the present disclosure propose a self-attention feature extraction module, which can simultaneously complete the extraction of channel attention features and the mining of location attention information. The extraction of channel attention features includes operations such as adaptive average pooling and basic convolutional blocks; the mining of location attention information includes matrix transposition and softmax normalization operations. This module can efficiently extract the basic feature information of the backbone network generated for the concerned region.
[0137] 2. The present disclosure proposes a feature slice fusion module, which slices the features multiple times and fuses them with the high-level feature map using periodic interpolation, can effectively fuse the feature information from different layers, maximally retain semantic information and suppress noise. This module can fuse high-level features and low-level features at the same time, and can improve the segmentation accuracy of abnormal regions in the image.
[0138] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented. The computer-readable storage medium can be a non-volatile computer-readable storage medium.
[0139] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to perform the above method.
[0140] The electronic device may be provided as a terminal, a server, or a device in other forms.
[0141] Figure 9 FIG. 4 is a block diagram of an electronic device 800 according to an embodiment of the present disclosure. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, or the like.
[0142] Refer to Figure 9 , the electronic device 800 may include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.
[0143] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.
[0144] The memory 804 is configured to store various types of data to support the operation of the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, and the like. The memory 804 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0145] The power component 806 provides power to various components of the electronic device 800. The power component 806 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the electronic device 800.
[0146] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0147] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 further includes a speaker for outputting audio signals.
[0148] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.
[0149] The sensor component 814 includes one or more sensors for providing an assessment of the various aspects of the state of the electronic device 800. For example, the sensor component 814 can detect the on / off state of the electronic device 800, the relative positioning of components, such as the display and the keypad of the electronic device 800. The sensor component 814 can also detect a change in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor component 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 814 can also include a light sensor, such as a CMOS or a CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 814 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0150] The communication component 816 is configured to facilitate communication between the electronic device 800 and other devices in a wired or wireless manner. The electronic device 800 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0151] In an exemplary embodiment, the electronic device 800 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0152] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 804 including computer program instructions, and the above computer program instructions can be executed by a processor 820 of the electronic device 800 to complete the above method.
[0153] Figure 10 is a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 can include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0154] The electronic device 1900 can further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an Input / Output (I / O) interface 1958. The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0155] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.
[0156] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0157] A computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but is not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punch card or raised structures in a groove having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0158] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0159] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0160] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0161] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0162] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0163] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0164] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of the terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. An attention-based image processing method, characterized in that, Including: Performing hierarchical multi-level feature extraction on the input image to respectively obtain a plurality of corresponding basic feature maps, where the basic feature maps include low-level features and high-level features; Successively performing self-attention feature extraction and feature enhancement processing on the basic feature maps to respectively obtain enhanced feature maps corresponding to the basic feature maps; Performing feature slice fusion processing on the enhanced feature map of the high-level features to obtain a high-level attention coefficient and a high-level segmentation feature map; Performing feature guidance processing on the enhanced feature map of the low-level features using the high-level attention coefficient to obtain a guidance fusion feature map; Performing anomaly detection using the high-level segmentation feature map and the guidance fusion feature map to obtain an anomaly detection result.
2. The method according to claim 1, wherein The successively performing self-attention feature extraction and feature enhancement processing on the basic feature maps to respectively obtain enhanced feature maps corresponding to the basic feature maps includes: Performing channel attention feature extraction and location attention information mining on the basic feature maps respectively to correspondingly obtain channel attention features and location attention information; Performing element-wise addition processing on the channel attention features and the location attention information to obtain a primary feature map; Performing a feature enhancement operation on the primary feature map to obtain an enhanced feature map corresponding to the basic feature map.
3. The method according to claim 1, wherein Wherein, The high-level features include front-layer features, middle-layer features, and back-layer features; The performing feature slice fusion processing on the enhanced feature map of the high-level features to obtain a high-level attention coefficient and a high-level segmentation feature map includes: Performing an upsampling operation on the enhanced feature map of the middle-layer features and performing an element-wise addition operation with the enhanced feature map of the front-layer features to obtain a high-level fusion feature map; Performing slice fusion processing on the high-level fusion feature map using the enhanced feature map of the back-layer features to obtain the high-level attention coefficient and the high-level segmentation feature map.
4. The method according to claim 3, wherein The performing slice fusion processing on the high-level fusion feature map using the enhanced feature map of the back-layer features of the high-level features to obtain the high-level attention coefficient and the high-level segmentation feature map includes: Performing slice processing on the high-level fusion feature map to obtain a plurality of slice layers; Inserting the enhanced feature map of the back-layer features between two adjacent slice layers and performing channel merging processing to obtain a first merged feature; Performing basic convolution processing on the first merged feature and adding the result of the basic convolution processing to the high-level fusion feature to obtain a second merged feature; Performing the basic convolution processing based on the second merged feature to obtain a third merged feature; Obtaining the high-level attention coefficient and the high-level segmentation feature map using the second merged feature and the third merged feature.
5. The method according to claim 1, characterized in that The performing feature guidance processing on the enhanced feature map of the low-level features using the high-level attention coefficient to obtain a guidance fusion feature map includes: Performing an element-wise multiplication operation between the high-level attention coefficient and the enhanced feature map of the low-level features respectively to obtain a refined feature map corresponding to the low-level features; Performing an element-wise addition operation on the enhanced feature map of the low-level features and the refined feature map respectively to obtain an aggregated feature map; The guiding fusion feature map is obtained by adding the aggregation features corresponding to the low-level features respectively.
6. The method according to claim 1, characterized in that, Performing anomaly detection by using the high-level segmentation feature map and the guiding fusion feature map, the anomaly detection result includes: Performing the feature slice fusion process by using the high-level segmentation feature map and the guiding fusion feature to obtain an image prediction feature; Performing anomaly detection on the image feature to obtain the anomaly detection result.
7. The method according to claim 6, wherein Performing anomaly detection on the image feature to obtain the anomaly detection result, includes: Using a decision tree model, taking the image feature as the root node to perform anomaly region detection; Judging and performing anomaly type detection based on the detected anomaly region; Generating an anomaly detection result based on the detected anomaly region and anomaly type.
8. An attention-based image processing device, characterized in that, The device includes: A basic feature map extraction module, configured to perform hierarchical multi-level feature extraction on the input image to obtain a plurality of corresponding basic feature maps respectively, where the basic feature map includes low-level features and deep features; A self-attention feature extraction module, configured to perform self-attention feature extraction and feature enhancement processing on the basic feature map in sequence to obtain an enhanced feature map corresponding to the basic feature map respectively; A feature slice fusion module, configured to perform a feature slice fusion process on the enhanced feature map of the deep feature to obtain a high-level attention coefficient and a high-level segmentation feature map; A guiding fusion module, configured to perform feature guiding processing on the enhanced feature map of the low-level feature by using the high-level attention coefficient to obtain a guiding fusion feature map; An anomaly detection module, configured to perform anomaly detection by using the high-level segmentation feature map and the guiding fusion feature map to obtain an anomaly detection result.
9. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is implemented.