A medical image organ segmentation method based on decoupled 3D self-attention network
By constructing a decoupled 3D self-attention network model and combining channel and spatial attention information, the problem of unclear effectiveness of the attention mechanism in existing medical image organ segmentation is solved, and more efficient and flexible medical image organ segmentation is achieved.
Patent Information
- Application Number
- CN202311001756.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-10
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-08-10
AI Technical Summary
In existing medical image organ segmentation technologies, the effectiveness of the attention mechanism has not been fully understood, and existing methods have limitations in computational cost and adaptability, which affects the application effect of medical image segmentation network models.
A decoupled 3D self-attention network model is constructed, and 3D self-attention processing of medical images is performed through the DSM module. By combining channel and spatial attention information, the effectiveness of the attention mechanism is explained using the NMF theory. A decoupled self-attention module DSM is designed and inserted into the medical image semantic segmentation network to reduce computational costs and improve segmentation performance.
It effectively reduces the computational cost, improves the accuracy and flexibility of medical image organ segmentation, enhances the performance of the network model, makes it more adaptable, and can better complete the task of medical image organ segmentation.
Smart Images

Figure CN117036704B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image organ segmentation, and in particular to a medical image organ segmentation method based on a decoupled 3D self-attention network. Background Art
[0002] Medical image organ segmentation technology is gaining increasing attention in the field of organ lesion diagnosis because it can quickly identify organ regions, providing important support for doctors' clinical diagnosis of organ lesions. The rapid advancement of medical image organ segmentation technology is due to the continuous development of medical image segmentation network models. The attention mechanism has now become an essential module of medical image segmentation network models. However, the reason why the attention mechanism is effective has been an unexplored question. The attention mechanism is generally believed to focus on important feature information and suppress unimportant feature information, but researchers have found that the consistency between the size of the attention weight and feature importance is weak. Exploring the fundamental reasons for the effectiveness of the attention mechanism can reveal the basis for the decisions made by medical image segmentation network models when performing organ segmentation, which can provide effective clues for clinicians to diagnose organ lesions.
[0003] The working process of channel attention in the SENet network is as follows: first, the input layer features C×H×W are globally averaged pooled channel by channel (i.e., the Squeeze step), and the global channel feature information C×1×1 is extracted. Then, the correlation between the global channel features is further explored through two fully connected layers (i.e., the Excitation step). Finally, Sigmoid is introduced to obtain weights and reweight the input layer features (i.e., the Reweight step). FMMNet believes that the purpose of the attention mechanism is to apply a feature map multiplication to the input layer features. This operation converts the linear piecewise function in the network model into a high-order piecewise function, thereby improving the performance of the network model. Compared with the attention mechanism, the self-attention or Transformer model obtains the self-attention map by calculating the correlation between the features of any point in the feature map and all the features of the points, thereby capturing long-distance dependencies. The HamNet network believes that the global modeling ability of self-attention is related to the low-rank embedding characteristics of the input features. For example, a large matrix is decomposed into a low-rank matrix through non-negative matrix decomposition (NMF), and then the large matrix is reconstructed through the coefficient matrix. However, FMMNet suffers from gradient explosion when stacking too many feature map multiplication layers. HamNet cannot preserve network weights like deep neural networks (DNNs) when using NMF for matrix decomposition. It is also sensitive to random seeds during testing. Therefore, the above methods greatly limit their adaptability and flexibility in real-world application scenarios. Summary of the Invention
[0004] In order to solve the problem that the above-mentioned existing technologies cannot effectively perform medical image organ segmentation, the present invention provides a medical image organ segmentation method based on a decoupled 3D self-attention network, which can provide a decoupled 3D self-attention network model to better complete the medical image organ segmentation task.
[0005] In order to achieve the above technical objectives, the present invention provides the following technical solutions: a medical image organ segmentation method based on a decoupled 3D self-attention network, comprising:
[0006] Constructing a decoupled 3D self-attention network model, wherein the decoupled 3D self-attention network model includes a medical image semantic segmentation network and a DSM module, wherein the DSM module is inserted into the backbone network of the medical image semantic segmentation network, and the DSM module performs 3D self-attention processing on the features of the medical image;
[0007] Obtain medical images, segment them using the decoupled 3D self-attention network model, and generate organ segmentation results for the medical images.
[0008] Optionally, the process of performing 3D self-attention processing on the medical image using the DSM model includes:
[0009] Acquiring channel attention information and spatial attention information according to input features of the DSM model; combining the channel attention information and the spatial attention information; and fusing the combined attention information;
[0010] Based on the fused attention information, self-attention information is obtained; based on the self-attention information, the input features are corrected and aggregated to generate output features.
[0011] Optionally, the process of obtaining channel attention information includes:
[0012] C=σ(Conv 1×1 (Conv 1×1 (X)×Softmax(Conv 1×1 (X))))
[0013] Among them, σ represents the Sigmoid function, C is the channel attention information, X is the input feature, Conv 1×1 Indicates a 1×1 convolution operation, and Softmax indicates a Softmax function operation.
[0014] Optionally, the process of obtaining spatial attention information includes:
[0015] S=σ(Conv 1×1 (X)×Softmax(MaxPooling(Conv 1×1(X)))
[0016] Among them, S is the spatial attention information, and MaxPooling means performing the maximum pooling operation.
[0017] Optionally, the channel attention information and the spatial attention information are combined in parallel and in series respectively.
[0018] Optionally, the combined attention information is fused channel by channel by setting weights for the parallel combination results and the serial combination results respectively, wherein the weights of the parallel combination results and the serial combination results are obtained by selecting kernel convolution processing.
[0019] Optionally, the process of obtaining self-attention information includes:
[0020] The fused attention information is processed by the Softmax function to generate self-attention information.
[0021] Optionally, the output feature generation process includes:
[0022]
[0023] Among them, Z is the output feature, Represents the spatial aggregation operation on each channel, M is the self-attention information, X is the input feature, represents broadcast element-by-element addition, X' is the 3D input feature, x'=σ(Conv 1×1 (x)), ⊙ is the matrix multiplication operation, σ represents the Sigmoid function, Conv 1×1 Indicates a 1×1 convolution operation.
[0024] The present invention has the following technical effects:
[0025] Through analysis, this paper discovered that the attention mechanism in ADNN corresponds to the basis matrix correction in NMF, and that the update of weight parameters in ADNN is consistent with the coefficient matrix correction in NMF. Based on this discovery, the paper aims to explain the fundamental reasons why the attention mechanism is effective and designs a decoupled 3D self-attention network model to better perform medical image organ segmentation tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] Figure 1 A schematic diagram of a process for alternately updating attention and network weight parameters between network layers provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of a process for iteratively correcting a basis matrix N and a coefficient matrix H in NMF according to an embodiment of the present invention;
[0029] Figure 3 Schematic diagram of the decoupled 3D self-attention module DSM proposed in the present invention provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0031] The following is an analysis of existing related technologies:
[0032] 1. Attention mechanisms in deep neural networks. Mainstream methods include channel attention, spatial attention, and hybrid attention. Hybrid attention combines channel attention and spatial attention in parallel or series to obtain the final attention. The decoupled 3D self-attention network of the present invention is dedicated to creating a 3D attention, but in order to effectively reduce computational costs, the 3D attention is first decoupled into channel and spatial attention, and then recoupled back to the 3D shape;
[0033] 2. Self-attention mechanism in deep neural networks. The self-attention mechanism can be considered as a special type of attention mechanism. The NLNet network obtains the self-attention map by calculating the correlation between any query point and all pixels in the feature map. The purpose is to model global context information and capture long-distance dependencies. However, the self-attention maps obtained by different query points in NLNet are the same. Based on this discovery, GCNet developed a simple SENet-like NLNet network, which greatly reduced the computational burden and complexity. Based on the shared self-attention map of different query points, the present invention can convert 3D attention into 3D self-attention by simply applying a Softmax operation, in order to alleviate the burden generated by the 3D self-attention calculation;
[0034] 3. Theoretical support for the attention mechanism. Traditional methods of explaining the effectiveness of the attention mechanism often use attention visualization or importance measurement. However, studies have found that there is no direct connection between the attention map and the importance of features. FMMNet believes that the attention mechanism is a feature map multiplication applied to the input layer features. This operation converts the linear piecewise function into a high-order piecewise function, which makes the attention mechanism effective. HamNet believes that the reason why the self-attention mechanism is effective is that it can effectively mine the low-rank embedding of the input features, and the original large matrix can be reconstructed by linear combination of the low-rank embeddings. Since FMMNet will experience gradient explosion when stacking too many feature map multiplication layers, HamNet cannot save network weights like deep neural networks (DNNs) when using NMF for matrix decomposition. At the same time, it is also sensitive to random seeds during testing. Therefore, the present invention believes that the essential reason why the attention mechanism is effective is that the attention mechanism improves the performance of the network model by correcting the input features.
[0035] Based on the above analysis results, this application conducts relevant analysis and design:
[0036] like Figure 1 As described above, the present invention analyzes the correspondence between the attention mechanism and weight parameter update in ADNN and the correction of the basis matrix and coefficient matrix in NMF, and on this basis concludes that the essential reason why the attention mechanism is effective is that it can correct the input features.
[0037] When optimizing between any two layers, a DNN only updates the network weight parameters. However, an attention-based deep neural network (ADNN) not only updates the weight parameters between any two layers but also reweights the features of the previous layer (input layer). When performing non-negative matrix factorization, NMF decomposes a large matrix into two low-rank matrices: a basis matrix and a coefficient matrix. During optimization, the basis matrix is first corrected, followed by the coefficient matrix, and the large matrix is approximated through continuous iteration.
[0038] 1) The connection between ADNN and NMF
[0039] like Figure 1 As shown, the present invention uses any two consecutive layers F in DNN l and F l+1 For example, in the process of network error back propagation, only the network weight W is updated l , but the input layer feature F l However, in ADNN, the present invention considers that the distance between any two consecutive layers is t i Use attention to update F at all times l , at t i+1 Update the weights by back propagating the error at all times. By alternating so that F′l+1 faster convergence to F l+1 . As Figure 2 shown, NMF approximately decouples the large matrix V into a low-rank basis matrix N and a coefficient matrix H, and iteratively corrects the basis matrix N and the coefficient matrix H by measuring the difference between V and V' = NH. The specific correction process is that, at time t i , the correction parameter of the basis matrix N is calculated to correct the basis matrix N, and at time t i+1 , the correction parameter of the coefficient matrix H is calculated to correct the basis matrix H.
[0040] 2) Explain the essential reason why the attention mechanism is effective using NMF theory
[0041] NMF approximately decouples the large matrix V into a low-rank basis matrix N and a coefficient matrix H, and iteratively corrects the basis matrix N and the coefficient matrix H by measuring the difference between V and V' = NH. The present application can use a stochastic gradient descent algorithm to update the basis matrix N and the coefficient matrix H respectively, as shown in formulas (1) and (2),
[0042] N' = N - a1(-VH T +NHH T ), (1)
[0043] H' = H - a2(-N T V+N T NH), (2)
[0044] wherein a1 and a2 represent learning rates, the superscript T represents matrix transposition, N' is the updated basis matrix, and H' is the updated coefficient matrix. The present application can set a1 = N / (NHH T ) and a2 = H / (N T HH), and after substituting formulas (1) and (2), formulas (3) and (4) are obtained respectively,
[0045] N' = [(VH T ) / (NHH T )] o N, (3)
[0046] H' = [(N T V) / (N T NH)] o H, (4)
[0047] wherein o represents matrix point multiplication, i.e. element-wise multiplication, and through formula (3) it can be concluded that NMF corrects the basis matrix N through the basis matrix correction factor (VH T ) / (NHH T ), and corrects the coefficient matrix H through the coefficient matrix correction factor.
[0048] In DNN, the weight parameter W of the lth layer network is updated by stochastic gradient descent l , as shown in formula (5),
[0049] W′ l =W l -η(-F l T F l+1 +F l T F l W l ), (5)
[0050] Among them, the update weight W in formula (5) l Corresponding to the correction of the coefficient matrix H in NMF, let η = W l / (F l T F l W l ), then it is converted into weight W after substituting it into formula (5) l The correction form of is shown in formula (6):
[0051] W′ l =[(F l T F l+1 ) / (F l T F l W l )]⊙W l. (6)
[0052] In DNN, only the weight W is updated l , but for the input feature F l No changes were made. With the help of NMF theory, the present invention speculates that ADNN uses the attention mechanism to focus on the input feature F l Correction is performed, as shown in formula (7):
[0053] F′ l =[(F l+1 W l T ) / (F l W l W l T )]⊙F l. (7)
[0054] Although the present invention can convert the correction form in formula (7) into the input feature F l The parameter update form is as shown in formula (8), but the obtained δ l+1 W lT It is very difficult because local convolution kernels are usually used in DNN, so it is impossible to obtain a large matrix W. l ,
[0055] F′ l =F l -λ(-F l+1 W l T +F l W l W l T )=F l -λ(δ l+1 W l T ), (8)
[0056] Among them, δ l+1 =-F l+1 +F l W l .
[0057] Based on the above discussion, this paper designs a decoupled self-attention module DSM to simulate the input feature F l The correction factor M is regarded as attention and is expressed as M≈(F l+1 W l T ) / (F l W l W l T ). Based on the correction factor M, the present invention can re-express the role of attention as the input feature F l The correction of is shown in formula (9):
[0058] F′ l =Attention⊙F l =M⊙F l . (9)
[0059] 2. The present invention designs a decoupled self-attention module (DSM) to obtain an attention map by simulating a correction factor. By inserting this module into the backbone of a classic medical image semantic segmentation network, the decoupled 3D self-attention network (DSNet) proposed in the present invention can be obtained.
[0060] The core content of the decoupled 3D self-attention network DSNet proposed in this paper is as follows:
[0061] 1) In the DSM module, the channel attention C and spatial attention S are first obtained separately, and then the channel attention C and spatial attention S are combined in parallel and series to recouple back to the original 3D shape;
[0062] 2) Parallel and series have their own advantages and disadvantages, the application sets a selection kernel convolution module to automatically determine the fusion weight of parallel and series according to the input data to obtain 3D attention A;
[0063] 3) 3D self-attention M can be obtained by applying Softmax operation to 3D attention A, then the corrected feature is aggregated, and finally, the corrected feature is added to the input feature F by broadcasting in a way of element-wise addition, that is, the 3D self-attention process is completed; l l 4) The DSM module is inserted into the classic medical image semantic segmentation network backbone, and the decoupled 3D self-attention network DSNet proposed in the application is obtained.
[0064] 4) The DSM module is inserted into the classic medical image semantic segmentation network backbone, and the decoupled 3D self-attention network DSNet proposed in the application is obtained.
[0065] The application designs a decoupled 3D self-attention network DSNet applied to medical image organ segmentation.
[0066] 1) First, the decoupled 3D attention in the DSM module is obtained,
[0067] In order to be more extensive and representative, the application uses X to represent the input feature (corresponding to the input feature F l described above), since the input X is 3D, that is, CxHxW, then the obtained attention should also be 3D. However, 3D attention brings huge computational cost, therefore, the application first decouples the 3D attention into channel attention C and spatial attention S. As shown in the following formula, the channel attention represents Figure 3 The process of generating channel attention C can be represented as C = σ (Conv 1×1 (Conv 1×1 (X) x Softmax (Conv 1×1 (X))), where σ represents the Sigmoid function; the spatial attention represents The process of generating channel attention S can be represented as S = σ (Conv 1×1 (X) x Softmax (MaxPooling (Conv 1×1 (X))), where σ represents the Sigmoid function;
[0068] 2) Secondly, the 3D attention shape is coupled back,
[0069] Channel attention C and spatial attention S are generally coupled back to 3D attention A through parallel and series, for parallel, the essence is C and S addition, that is, P = C + S, then series is C and S multiplication, that is, Q = C x S, However, both parallel or series have their own advantages, so the present application will select kernel convolution module into the DSM, by adaptive adjustment of the fusion weight of P and Q to obtain 3D attention A, that is Where α = Conv SK (Q, P), Conv SK Representing the selection kernel convolution, Representing the channel-wise addition; "selection kernel convolution" itself is a convolution module that has been proposed (from SKNet network model), through which the proportion of two different convolution kernels can be adaptively determined and weighted combined into a new convolution kernel.
[0070] 3) Then, obtain the decoupled 3D self-attention in the DSM module,
[0071] The computational complexity of the calculation of the conventional method of self-attention reaches Inspired by GCNet network, the present application can obtain 3D self-attention M by directly applying Softmax operation to 3D attention A, that is M = Softmax(A). After obtaining 3D self-attention M, M is applied to 3D input feature x' as a correction factor to complete the feature correction process, where x' = σ(Conv 1×1 (X)), then perform feature aggregation operation and add to the original input feature X pixel by pixel to obtain output feature Z, that is Where Representing the spatial aggregation operation on each channel, Indicates broadcast element-wise addition;
[0072] 4) Finally, after inserting into the second convolution layer of the last stage of the classic medical image semantic segmentation network backbone, the decoupled 3D self-attention network DSNet is obtained,
[0073] For example, the present application takes the classic ResNet50 network backbone as an example, uses FCN as the segmentation head, and obtains the decoupled 3D self-attention network DSNet designed by the present application. The specific network structure of the network is shown in Table 1 below.
[0074] Table 1
[0075]
[0076]
[0077]
[0078]
[0079]
[0080] 1.Input layer features are often 3D, but existing attention mechanisms obtain 2D attention maps, but attention should be pixel by pixel, so 3D attention mechanism is a more suitable solution.
[0081] The present application is dedicated to building 3D attention to correct input features, but in order to effectively reduce the computational cost, first decouple the 3D attention to channel and spatial attention, and then recouple back to 3D shape, greatly reducing the computational cost while obtaining 3D attention;
[0082] 2.What is the relationship between 3D attention mechanism and 3D self-attention mechanism needs to be explored, which is also an important factor restricting the development of medical image segmentation network model.
[0083] Because the 3D self-attention mechanism has too high computational complexity, if the relationship between the two can be explored, the 3D self-attention can be directly converted from the 3D attention mechanism, which can effectively solve the problem. Based on the sharing of different query points, the 3D attention is converted into 3D self-attention shape by applying Softmax operation, thereby greatly reducing the overhead of calculating 3D self-attention;
[0084] 3.The attention network module constructed under the existing attention mechanism or self-attention mechanism explanation theory lacks adaptability and flexibility, and it is a better solution to build an attention network module that can better incorporate DNN.
[0085] For example, FMMNet will have gradient explosion when stacking too many feature graph multiplication layers, HamNet cannot save network weights like deep neural network (DNN) when using NMF for matrix decomposition, and is also sensitive to random seeds during testing. The present application found through analysis that the attention mechanism in ADNN corresponds to the base matrix correction in NMF, and the update of the weight parameters in ADNN is consistent with the coefficient matrix correction in NMF. Therefore, the present application believes that the reason why the attention mechanism is effective is that the attention mechanism corrects the input features to improve the performance of the network model. Based on the above analysis, a decoupled 3D self-attention network model is designed, which can fully utilize the advantages of the attention mechanism to better complete the medical image organ segmentation task.
[0086] The above is only the preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any skilled person in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A medical image organ segmentation method based on decoupled 3D self-attention network, characterized by: include: Constructing a decoupled 3D self-attention network model, wherein the decoupled 3D self-attention network model includes a medical image semantic segmentation network and a DSM module, wherein the DSM module is inserted into the backbone network of the medical image semantic segmentation network, and the DSM module performs 3D self-attention processing on the features of the medical image; Obtain medical images, segment them using the decoupled 3D self-attention network model, and generate organ segmentation results for the medical images; The process of performing 3D self-attention processing on medical images through the DSM module includes: According to the input features of the DSM module, channel attention information and spatial attention information are obtained; the channel attention information and the spatial attention information are combined; and the combined attention information is fused; Based on the fused attention information, self-attention information is obtained; based on the self-attention information, the input features are corrected and aggregated to generate output features; The process of obtaining channel attention information includes: in, represents the Sigmoid function, C is the channel attention information, X is the input feature, Conv 1×1 Indicates a 1×1 convolution operation, and Softmax indicates a Softmax function operation. The process of acquiring spatial attention information includes: Among them, S is the spatial attention information, and MaxPooling means performing the maximum pooling operation; Combining the channel attention information and the spatial attention information in parallel and in series respectively; The combined attention information is fused channel by channel by setting weights for the parallel combination results and the serial combination results respectively, where the weights are obtained by performing kernel convolution processing on the parallel combination results and the serial combination results; The process of acquiring self-attention information includes: Performing Softmax function processing on the fused attention information to generate self-attention information; The process of generating output features includes: Among them, Z is the output feature, Represents the spatial aggregation operation on each channel, M is the self-attention information, X is the input feature, represents broadcast element-wise addition, X' is the 3D input feature, , ☉ is the matrix dot multiplication operation, Represents the Sigmoid function, Conv 1×1 Indicates a 1×1 convolution operation.
Citation Information
Patent Citations
Lane line detection system based on geometric attention perception
CN111582201A
Medical image processing method based on attention mechanism and related equipment
CN114120030A