A method, device, and storage medium for detecting a prostate lesion area
Through dynamic multi-scale perception adaptive integration network, combined with dynamic local pooling and global attention mechanism, the problems of size changes and blurred boundaries in prostate lesion area detection are solved, and high-quality lesion area detection and segmentation are achieved.
Patent Information
- Application Number
- CN202310486436.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-28
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-04-28
AI Technical Summary
The existing prostate lesion area detection methods are difficult to adapt to prostate size changes and blurred boundaries, and cannot fully learn and dig up hidden information in medical images, resulting in poor segmentation effect.
Adaptive integration network based on dynamic multi-scale perception is adopted, and the adaptive integration of multi-scale information is optimized through dynamic local pooling and global efficient attention mechanisms, combined with dual-flow attention and residual attention mechanisms, and the adaptive integration of multi-scale information is achieved.
It improves the detection accuracy and segmentation effect of prostate lesion areas, can extract high-quality information at different scales, adapt to changes in prostate images, enhances the network's attention to foreground information and suppresses background noise.
Smart Images

Figure CN116542924B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and specifically to a method, device, and storage medium for detecting prostate lesion regions based on a dynamic multi-scale perception adaptive integration network. Background Art
[0002] Prostate segmentation refers to detecting the prostate region in an input prostate MR image. If it is prostate lesion segmentation, the lesions are segmented from the original image if any. Since medical images are very different from conventional natural images, directly applying conventional segmentation algorithms to medical image segmentation will inevitably result in performance degradation. This is especially true for the segmentation of prostate lesion regions. Accurately segmenting the prostate and lesion regions from magnetic resonance (MR) images is crucial for the diagnosis and treatment planning of prostate diseases. Especially for common prostate diseases such as prostate cancer, prostatitis, and benign prostatic hyperplasia, in the general diagnosis of diseases, medical images are usually segmented manually by several experts, which is a very time-consuming process. With the development of deep learning, convolutional neural networks have made great progress in the field of prostate and lesion segmentation, and thus more branch studies on prostate and prostate lesion region segmentation have emerged. These studies have played a very important role both in the theoretical research on the prostate and its diseases and in front-line medical practice.
[0003] According to different input forms, prostate and prostate lesion segmentation can be divided into two categories: 2D prostate and its lesion region detection and 3D prostate and its lesion region detection. Among them, the input of 2D prostate and its lesion region detection is a single MR image, which is a grayscale image with 1 channel. The input of 3D prostate and its lesion region detection is a series of consecutive MR images. This is because for a patient, the imaging of the prostate and its lesion region is a series of consecutive images, that is, different slices in the axial direction. Since there is a connection between the slices of the same patient, 3D prostate and its lesion region detection has emerged. With the progress of technology and hardware, recently, prostate and its lesion region detection with video input has also emerged, that is, using more input information to achieve more accurate segmentation results.
[0004] However, the current state-of-the-art automatic MR image segmentation methods face some challenges. There is a lack of clear boundaries and high contrast between the prostate and surrounding tissues, making it difficult to accurately extract the prostate from the background. In addition, the complexity of the background texture and the large variations in the scale, shape, and intensity distribution of the prostate itself make the segmentation more complex. Moreover, the characteristics of the prostate lesion area are visually more difficult to distinguish compared to the prostate area, and professional doctors also need to view data in multiple modalities to determine the location of the lesion area. At the same time, due to the scarcity of information in medical images, multiple methods must be used for multi-dimensional information extraction to enable the network to make up for the innate lack of information in the images with rich information extracted in multiple ways. However, the current methods cannot fully learn and exploit the hidden information in the images. Summary of the Invention
[0005] Aiming at the problem that the current prostate lesion area detection methods still use conventional fixed-parameter layers to infer the input prostate images, which are difficult to adapt to the changes in prostate size and blurred boundaries in the input images, the present invention provides a prostate lesion area detection method based on a dynamic multi-scale perception adaptive integration network. The method uses the input prostate MR images for lesion area detection and optimizes and updates through a network of dynamic local pooling and global efficient attention, achieving high-quality lesion area detection in the scenario of a given prostate image.
[0006] To this end, the present invention provides the following technical solutions:
[0007] A prostate lesion area detection method based on a dynamic multi-scale perception adaptive integration network, comprising the following steps:
[0008] A. Obtain prostate input images from the prostate dataset and obtain tensors;
[0009] B. Input the tensors into a feature encoder to obtain multi-scale encoded features based on each image through the feature encoder;
[0010] C. For the encoded features, obtain a richer feature representation through the corresponding feature enhancement layer
[0011] D. Perform feature decoding on the richer feature representation through a decoder to obtain the final prostate segmentation prediction result, including:
[0012] D1. Through a multi-level integration module with convolution, use the characteristic of hierarchical complementarity to establish a multi-level feature pyramid, and adaptively encode the multi-scale information of an image into the current pyramid feature vector to obtain the feature information of different levels integrated globally and locally;
[0013] The mechanism of hierarchical complementarity includes: The features extracted by the network include local features and global features. Among them, the local features are extracted by the shallow convolutional network, which mainly contains information about the details and textures of the image, regardless of whether it belongs to the foreground or background; while the global features are mainly extracted by the deep convolutional module and the attention mechanism. The global features mainly include the part about high-level semantic information, such as the position, the difference between foreground and background, etc. It is precisely because the focuses of these two types of features are local and global respectively, and for prostate lesion segmentation, the position and detail information are equally important. Therefore, when integrating, the complementary characteristics of global and local features between different levels can be utilized to obtain better prediction results.
[0014] The integration module includes: Through convolution and fusion operations, obtain the integrated feature maps of two consecutive layers; then perform channel transformation and normalization operations on the obtained integrated feature maps to obtain the output of a certain layer of the integration module. Among them, convolution can use multiple convolution kernels of different sizes to fuse neighborhood information or transform the number of channels for better fusion in the next step.
[0015] D2. Dynamically fuse the feature layers between different levels in a progressive manner in multiple stages to obtain a more accurate and richer fused feature representation. And the fusion weights are automatically learned through convolution operations. By adaptively dynamically weighting and fusing the features between different levels, the final prostate segmentation prediction result map is obtained.
[0016] Furthermore, step A includes:
[0017] Divide the prostate dataset into a training set and a test set with a fixed number. Both the training set and the test set should have two subsets, namely the input image and the ground truth.
[0018] First, perform data augmentation on the data in the prostate training set, including but not limited to resizing the input prostate image to H×W, and the resizing size can be set during use to best match the network model; secondly, use random flipping and random cropping with random probabilities; perform format conversion on the augmented image to convert it into a tensor that the network can process, thus obtaining a tensor of batchsize size. It should be noted that the same operations should also be performed on the ground truth in the prostate training set.
[0019] For the data in the prostate test set, the operations are different from the above. First, the size of the input image will be adjusted, and then the image will be directly tensorized. After obtaining a tensor of batchsize size, it will be directly fed into the network for testing.
[0020] Furthermore, the batch size is set to 8; H×W is set to 224×224.
[0021] Furthermore, the feature encoder adopts the ResNet architecture, and the last two layers are discarded to retain the spatial structure. Then, a global-local complementary module is added after the output of each layer to extract multi-scale context information. And the features of all layers except the first layer are stored in the feature pyramid to facilitate the operation of subsequent modules. That is, the feature encoder generates 1 feature pyramid for each image, which includes 4 feature maps with different spatial resolutions and numbers of channels.
[0022] Furthermore, the ResNet architecture is the ResNet-50 architecture, in which the number of input channels of the first convolutional module is modified to 1 to adapt to the number of channels of the input image. The global-local complementary module is a module containing two branches, namely the dynamic local pooling branch and the global efficient attention branch.
[0023] Dynamic local pooling includes dynamically allocating a suitable combination of pooling layers of appropriate sizes to each layer according to the position of the layer in the entire network to better extract local information. That is, a combination with a large number and large convolutional kernel sizes is used in the lower layers of the network to adapt to the large-size lower-layer feature maps; while in the higher layers of the network, a combination with a small number and small convolutional kernel sizes is used to adapt to the small-size higher-layer feature maps and prevent information loss.
[0024] Global efficient attention includes directly performing an operation of comparing the similarities of the input feature maps to obtain a similarity weight map. If the dimension of the feature map of this layer is relatively high, it can be first dimension-reduced to save computing resources. Then, by combining the weight map with the input image, it is calculated which regions should be focused on by the network globally.
[0025] Furthermore, the convolutional kernels of dynamic local pooling are 1, 3, 5, 7, 9 in sequence. Among them, those with a size greater than 3 use depthwise separable convolutions instead of ordinary convolutions to reduce the number of parameters.
[0026] Furthermore, in step D1, if it is for neighborhood fusion, the size of the dynamic kernel K t is all 3×3; while if it is for channel number integration, the size is all 1×1.
[0027] Furthermore, step C includes:
[0028] In the corresponding feature fusion layer, the features corresponding to the scales of the feature encoder are respectively used as inputs;
[0029] For the features at this scale, the dual-stream attention mechanism and the residual-attention mechanism are respectively used to further optimize the feature representation after passing through the feature encoder;
[0030] For the feature representations with the same spatial resolution after different attention transformations, pixel-level summation is adopted to obtain a more abundant fused feature representation.
[0031] Furthermore, for the features at each scale, different attention mechanisms are used to enable the network to focus on different information, thereby improving the richness and accuracy of the feature representation, including:
[0032] For the dual-stream attention, it is to utilize the attention to emphasize the characteristics of different regions. The foreground regions that the network should focus on under normal circumstances are calculated using spatial attention and channel attention. Subsequently, the weight map of the foreground attention is normalized and inverted so that it focuses on the background, better finding the foreground information hidden in the background features. Then, the fused feature representation is obtained using pixel-level summation and convolution operations.
[0033] For the residual-attention mechanism, it is the combination of the residual module with self-attention and mutual-attention. That is, multi-branch residual convolution operations are performed on the input feature representation. And the information flow between the branches is realized through the combination of self-attention weights and mutual-attention weights, thereby obtaining a more abundant feature representation.
[0034] Furthermore, the number of multi-branches in the residual-attention mechanism is set to 4, and the weighted average between the self-attention weight map and the mutual-attention weight map is 1.
[0035] The above technical solutions provided by the present invention have the following beneficial effects:
[0036] The present invention proposes a method for detecting prostate lesion regions based on a dynamic multi-scale perception adaptive integration network, which takes into account the coherence between multi-scale information in the input image. First, encoded features at multiple scales are obtained for each image through a feature encoder. In the feature encoder, dynamic pooling convolution is first used to focus on local information of the image, and a global efficient attention mechanism is used in combination to focus on global information of the image. The two are integrated through pixel addition and convolution operations to extract effective information at different scales in the input image to the greatest extent. Secondly, in order to avoid the negative impact of noise on the prediction results during the information extraction process, information enhancement operations are performed on the feature maps at each scale using a combination of multiple attentions. On the one hand, a two-stream attention mechanism is adopted to enable the network to pay more attention to foreground information and suppress noise; on the other hand, residual attention is used to strengthen the connection between each branch while the network extracts and integrates information, and the complementary characteristics between each branch are used for more accurate prediction. Experimental results show that the method for detecting prostate lesion regions based on the dynamic multi-scale perception adaptive integration network proposed by the present invention has accurate prediction effects on prostate lesion region segmentation and prostate segmentation.
[0037] For the above reasons, the present invention can be widely promoted in the field of prostate and its lesion regions. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic diagram of the input prostate MR in the embodiment of the present invention;
[0040] Figure 2 It is a flowchart of the method for detecting prostate lesion regions based on the dynamic multi-scale perception adaptive integration network in the embodiment of the present invention;
[0041] Figure 3 It is a schematic diagram of the structure of the dynamic adaptive pooling module in the embodiment of the present invention;
[0042] Figure 4 It is a schematic diagram of the structure of the two-stream attention module in the embodiment of the present invention;
[0043] Figure 5 It is a schematic diagram of the structure of the internal attention module of the residual-attention module in the embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] See Figure 2 , which shows a flowchart of a method for detecting prostate lesion regions based on a dynamic multi-scale perception adaptive integration network in an embodiment of the present invention. The method includes the following steps:
[0046] A. Obtain a prostate input image from a prostate dataset and obtain a tensor;
[0047] In a specific implementation, step A specifically includes:
[0048] A1. Obtain a prostate image:
[0049] The schematic diagram of the input prostate MR is as shown in Figure 1 . The prostate dataset is divided into a training set and a test set according to a certain ratio. And in both the training set and the test set, there are two sub-datasets, namely the input image and the ground truth.
[0050] A2. For the input prostate image, obtain a tensor T with the number of channels being the size of batchsize
[0051] Perform data augmentation on the input image and the corresponding ground truth in the prostate training set. First, adopt a random cropping strategy with a scale of s and a ratio of r for the input MR original image (i.e., a single-channel grayscale image) and the GT image, and resize it to H×W (the image resolution adopted in this method is 224×224). Then, use random flipping with a random probability. Subsequently, convert the enhanced grayscale image into a tensor that can be processed by the network, and use a data loader to adjust the number of channels of the output tensor to the size of batchsize, so that the network can converge more quickly. The value of batchsize here is set to 8.
[0052] For the input prostate image and the corresponding ground truth in the prostate test set, perform an operation of resizing it to H×W (the image resolution adopted in this method is 224×224). Then, convert the resized prostate grayscale image into a tensor that can be processed by the network, and use a data loader to adjust the number of channels of the output tensor to the size of batchsize, so that the network can converge more quickly. The value of batchsize here is set to 1.
[0053] B. Input the tensor into the feature encoder, and obtain the encoded features at multiple scales based on each image through the feature encoder.
[0054] In a specific implementation, step B specifically includes:
[0055] B1. Input the obtained tensor I t into the feature encoder:
[0056] The adopted feature encoder is of ResNet-50 architecture, where the number of input channels of the first convolutional module is modified to 1 to adapt to the number of channels of the input image, and at the same time, the last fully connected layer is removed to adapt to the subsequent modules.
[0057] B2. Obtain the encoded features at multiple scales
[0058] The feature encoder will generate 5 multi-scale feature maps with different spatial resolutions and numbers of channels for each image, that is whose resolutions and numbers of channels (W×H×C) are respectively Since there is too much noise and less useful information in the first layer, in subsequent information enhancement, only the features of the 1st - 5th layers are operated on.
[0059] Among them, step B2 specifically includes:
[0060] Each layer contains a residual convolutional module and a two-branch global-local complementary module, and these two modules are combined in series.
[0061] The residual convolutional module is a convolutional module with residual connections. The two-branch global-local complementary module contains a path of dynamic adaptive local pooling and a path of global efficient attention. As Figure 3 shown, dynamic adaptive local pooling uses convolutional kernels of different sizes for convolution. Among them, for low-level features, the more convolutional numbers contained in the dynamic adaptive pooling module, the larger the range covered by its convolutional kernel, to make up for the lack of information at the low level; for high-level features, since the information is already very sufficient and more integration and transformation are needed, the size of its convolutional kernel is relatively small and the number is relatively small. The whole process can be expressed as:
[0062] layers(i)∈{RL(j)|j=0,1}∪{OL(k)|1≤k≤5 - i,k∈N +} (1)
[0063] Among them, layers(i) represents the feature map of the i-th layer. RL represents a necessary part for all feature layers containing the dynamic adaptive local pooling module. The RL part includes a global average pooling, an upsampling, and a convolution operation with a convolution kernel size of 1. The OL part represents that for different feature layers, the composition of the OL part is different, and the convolution kernel sizes and numbers are different (i.e., the value range of k). The size of the convolution kernel is equal to 2k - 1.
[0064] The global efficient attention module utilizes the concept of similarity, combines the self-attention mechanism, enhances the network's ability to focus on the global region, enables the network to better learn more information from the global perspective, and thus makes more accurate predictions. And residual connections are used to prevent information loss during network operation. The entire process can be represented as:
[0065]
[0066] Among them, x represents the input feature map, and Normalize is the normalization operation.
[0067] At the end of each layer, the feature maps of the two branches are integrated through pixel-level addition and convolution operations.
[0068] C. For the encoded features, richer feature representations are obtained through the corresponding feature enhancement layers
[0069] In specific implementation, step C specifically includes:
[0070] C1. The features of each scale are sent to the dual-stream attention module to enhance the foreground and suppress background noise:
[0071] Conventional attention can only focus on the foreground area and ignore the foreground information hidden in the background area. Therefore, in the dual-stream attention module of the embodiment of the present invention, the spatial attention mechanism and the channel attention mechanism are combined, so that the network can focus on the foreground area while also paying attention to the foreground information hidden in the background area, as Figure 4 shown. At the end of the dual-stream attention module, the features of the foreground branch and the background branch are integrated in the form of convolution to learn the information therein for more accurate segmentation.
[0072] C2. The features of each scale are sent to the residual-attention module to further integrate information in a multi-branch form, and the attention mechanism is used to enhance the correlation between multi-branches in the form of residuals to jointly highlight the foreground target.
[0073] As shown Figure 5 in the figure, there are a total of 4 branches in the residual-attention module, and each branch has 2 residual convolution modules. Between each branch, an attention module is used to strengthen the association. The attention module is divided into two parts. One part focuses on self-attention, which is used to generate a weight-weighted matrix about itself; the other part is cross-attention, which uses the similarity between the two branches to generate a weight-weighted matrix when the two branches exchange information. The output of each branch is obtained by weighted summation of the feature maps of adjacent branches and its own branch feature map through the matrix calculated above. Finally, the results of the 4 branches are integrated to obtain the output of the residual-attention module.
[0074] D. Through the decoder for richer feature representations perform feature decoding to obtain the final segmentation prediction result of the prostate lesion area.
[0075] The decoder includes an integration module and a prediction module. The decoder receives a feature pyramid composed of 4 different-scale feature maps from the previous module. The integration module of the decoder starts from the feature map with the smallest size and the highest number of layers (i.e., 7×7×2048). By gradually integrating the feature maps of two consecutive sizes, a final integrated feature map is obtained. The integration operation mainly includes two steps: feature map stacking in the channel direction and feature transformation. The final integrated feature map will be sent to a prediction module, and the prediction module will predict the integrated feature map by category according to the number of classifications required by the task.
[0076] The specific steps are as follows:
[0077] D1. Through a multi-level integration module with convolution, use the characteristic of hierarchical complementarity to establish a multi-level feature pyramid, and adaptively encode the multi-scale information of an image into the current pyramid feature vector to obtain different levels of integrated feature information including global and local. The mechanism of hierarchical complementarity includes: the features extracted by the network include local features and global features. Among them, the local features are extracted by the shallow convolutional network, which mainly contains information about the details and textures of the image, and these information do not distinguish whether they belong to the foreground or background; while the global features are mainly extracted by the deep convolutional module and the attention mechanism, and the global features mainly include the part of high-level semantic information, such as position, foreground-background difference, etc. It is precisely because the focuses of these two types of features are local and global respectively, and for prostate lesion segmentation, position and detail information are equally important. Therefore, the complementary characteristics of global and local features between different levels can be used during integration to obtain better prediction results.
[0078] The integration module includes: obtaining integrated feature maps of two consecutive layers through convolution and fusion operations; then performing channel transformation and normalization operations on the obtained integrated feature maps to obtain the output of a certain layer of the integration module. Among them, multiple convolution kernels of different sizes can be used for convolution to fuse neighborhood information or transform the number of channels, so as to better perform the next fusion operation.
[0079] D2. Dynamically fuse the feature layers between different levels in a progressive manner in multiple stages to obtain a more accurate and richer fused feature representation. The fusion weights are automatically learned through convolution operations. By adaptively and dynamically weighted-fusing the features between different levels, the final prostate segmentation prediction result map is obtained. E. Training and optimization of the dynamic context-aware filtering network:
[0080] This method can be generally divided into two stages: training and inference. During training, the tensors of the training set are used as inputs to obtain the trained network parameters; during the inference stage, the parameters saved in the training stage are used for testing to obtain the final saliency prediction result.
[0081] This embodiment of the invention is implemented under the Pytorch framework. During the training stage, the ADAMW optimizer is used, the learning rate is 1e -3 , and the weight decay coefficient is 5e -2 , and the batch size is 8. During training, the spatial resolution of the image is 224×224, but the model can be applied in a fully convolutional manner to any resolution during testing.
[0082] The prostate lesion region detection method based on the dynamic multi-scale perception adaptive integration network proposed in this embodiment of the invention adopts dynamic local pooling combined with a global attention mechanism to incorporate the context information in the input prostate MR image into the current feature matrix, obtaining a feature vector containing global information and local detail information, and adapting to the scale change of the target. Secondly, in order to avoid misleading the final saliency result, this invention adopts a variety of attention complementary perception fusion methods, and uses a variety of attentions to enhance the foreground and suppress the background noise for the feature maps generated at each scale. Experimental results show that the prostate lesion region detection method based on the dynamic multi-scale perception adaptive integration network proposed in this invention can obtain accurate prediction results for many scenarios with prostate size changes and low input image quality.
[0083] Corresponding to the prostate lesion region detection method in the above embodiment, this embodiment of the invention also provides a prostate lesion region detection device based on the dynamic multi-scale perception adaptive integration network, including:
[0084] A tensor unit for obtaining a prostate input image and obtaining a tensor according to the prostate cancer dataset;
[0085] An encoding unit, configured to input the tensor obtained by the tensor unit into a feature encoder. The feature encoder uses dynamic pooling convolution to focus on local information of an image, combines a global efficient attention mechanism to focus on global information of the image, and integrates the local information and the global information through operations of pixel summation and convolution to obtain multi-scale encoded features based on each image.
[0086] An enhancement unit, configured to obtain a feature representation for the encoded features obtained by the encoding unit through a corresponding feature enhancement layer.
[0087] A prediction unit, configured to perform feature decoding on the feature representation obtained by the enhancement unit through a decoder to obtain a final prediction result for prostate lesion area segmentation.
[0088] For the prostate lesion area detection device according to the embodiments of the present invention, since it corresponds to the prostate lesion area detection method in the above embodiments, the description is relatively simple. For relevant similarities, please refer to the description of the prostate lesion area detection method part in the above embodiments, and details are not described herein again.
[0089] An embodiment of the present invention further provides a computer-readable storage medium, which stores computer instructions for causing a computer to execute the prostate lesion area detection method based on the dynamic multi-scale perception adaptive integration network as described above.
[0090] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting prostate lesion regions based on a dynamic multi-scale perception adaptive integration network, characterized in that, It includes the following steps: Obtain the prostate input image from the prostate cancer dataset and get the tensor Input the tensor into the feature encoder. Through the feature encoder, use dynamic pooling convolution to focus on the local information of the image, combine the global efficient attention mechanism to focus on the global information of the image, and integrate the local information and the global information through pixel addition and convolution operations to obtain multi-scale encoded features based on each image; For the encoded features, obtain the feature representation through the corresponding feature enhancement layer, including: in the corresponding feature fusion layer, use the features corresponding to the scale of the feature encoder as the input respectively; For the features of this scale, use the dual-stream attention mechanism and the residual-attention mechanism respectively to further optimize the feature expression after passing through the feature encoder; for the feature expressions with the same spatial resolution after using different attention transformations respectively, use pixel-level addition to obtain a more rich feature representation after fusion; the dual-stream attention mechanism uses spatial attention and channel attention to calculate the foreground area that the network should focus on under normal circumstances, and then performs normalization and inversion operations on the weight map of the foreground attention to make it focus on the background, and then uses pixel-level addition and convolution operations to obtain the feature representation after fusion; the residual-attention mechanism uses the residual module, self-attention, and mutual-attention to perform multi-branch residual convolution operations on the input feature representation, and the information flow between each branch is realized through the combination of self-attention weights and mutual-attention weights; Perform feature decoding on the feature representation through the decoder to obtain the final prostate lesion area segmentation prediction result.
2. The prostate lesion area detection method based on the dynamic multi-scale perception adaptive integration network according to claim 1, wherein Obtain the prostate input image and get the tensor according to the prostate cancer dataset, including: Divide the prostate dataset into training sets and test sets with a fixed number, where both the training set and the test set have two subsets, namely the input image and the ground truth; Perform data augmentation on the data in the prostate training set; Perform format conversion on the augmented image to convert it into a tensor that the network can process, and obtain a tensor of batchsize size; For the data in the prostate test set, adjust the size of the input image and directly tensorize the image to obtain a tensor of batchsize size.
3. The prostate lesion area detection method based on a dynamic multi-scale perception adaptive integration network according to claim 1, characterized in that: The feature encoder is of ResNet architecture, and the last two layers are discarded to retain the spatial structure, and then a global-local complementary module is added after the output of each layer to extract multi-scale context information; and the features of all layers except the first layer are stored in the feature pyramid.
4. The video saliency detection method based on a dynamic context-aware filtering network according to claim 3, characterized in that: The global-local complementary module includes a dynamic local pooling branch and a global efficient attention branch; The dynamic local pooling includes, at each layer of the network, dynamically allocating a pool of appropriate size for this layer according to the position of this layer in the entire network to better extract local information; The global efficient attention includes directly performing a similarity comparison operation on the input feature map to obtain a similarity weight map; using the combination of this weight map and the input image to calculate the area that needs to be focused on globally.
5. The prostate lesion region detection method based on a dynamic multi-scale perception adaptive integration network according to claim 1, characterized in that: Feature decoding is performed on the feature representation through a decoder to obtain the final prostate lesion area segmentation prediction result, including: Through a multi-level integration module with convolution, a multi-level feature pyramid is established by utilizing the hierarchical complementary characteristics, and the multi-scale information of an image is adaptively encoded into the current pyramid feature vector to obtain the feature information after integration at different levels including global and local information. In multiple stages, the feature layers between different levels are adaptively and dynamically weighted and fused in a progressive manner to obtain the final prostate segmentation prediction result map, and the fusion weights are automatically learned through convolution operations.
6. The prostate lesion region detection method based on a dynamic multi-scale perception adaptive integration network according to claim 1, wherein: The number of multi-branches in the residual-attention mechanism is set to 4, and the weighted average between the self-attention weight map and the mutual-attention weight map is 1.
7. A prostate lesion region detection device based on a dynamic multi-scale perception adaptive integration network, characterized in that, Including: A tensor unit for obtaining a prostate input image from a dataset of prostate cancer and obtaining a tensor An encoding unit for inputting the tensor obtained by the tensor unit into a feature encoder. Through the feature encoder, local information of the image is focused by using dynamic pooling convolution, and global information of the image is focused by combining with a global efficient attention mechanism. The local information and global information are integrated through pixel addition and convolution operations to obtain the encoded features at multiple scales based on each image. An enhancement unit for obtaining a feature representation for the encoded features obtained by the encoding unit through a corresponding feature enhancement layer, including: in the corresponding feature fusion layer, features corresponding to the scale of the feature encoder are respectively used as inputs; for the features at this scale, the dual-stream attention mechanism and the residual-attention mechanism are respectively used to further optimize the feature expression after passing through the feature encoder; for the feature expressions with the same spatial resolution after different attention transformations, pixel-level addition is used to obtain a more abundant fused feature representation; in the dual-stream attention mechanism, spatial attention and channel attention are used to calculate the foreground area that the network should focus on under normal circumstances, and then the weight map of the foreground attention is normalized and inverted so that it focuses on the background, and then a fused feature representation is obtained through pixel-level addition and convolution operations; in the residual-attention mechanism, the input feature representation is subjected to multi-branch residual convolution operations using a residual module, self-attention, and mutual-attention, and the information flow between branches is realized through the combination of self-attention weights and mutual-attention weights. A prediction unit for performing feature decoding on the feature representation obtained by the enhancement unit through a decoder to obtain the final prostate lesion area segmentation prediction result.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to execute a method for detecting a prostate lesion area based on a dynamic multi-scale perception adaptive integration network according to any one of claims 1-6.
Citation Information
Patent Citations
Plant disease and pest identification method based on attention mechanism and multi-level convolution characteristics
CN110188635A
Image salient object segmentation method and apparatus based on reciprocal attention between foreground and background
US20200372660A1