A multi-task whole heart CT segmentation method based on pooling Transformer
By adopting 3D pooled Transformer and dual-task learning methods in the whole-cardiac CT segmentation task, the problems of insufficient global information capture and blurred boundaries are solved, and efficient and accurate whole-cardiac segmentation and boundary constraints are achieved.
Patent Information
- Application Number
- CN202310609850.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-26
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-05-26
AI Technical Summary
The prior art is difficult to capture the global information of the image in the full-cardiac segmentation task, resulting in the need to improve the segmentation performance, and the cardiac images are affected by noise and low contrast, and the boundary blur is not easy to distinguish.
A multi-task full-cardiac CT segmentation method based on pooled Transformer is proposed, using 3D pooled attention mechanism to replace the traditional multi-head self-attention mechanism, reducing the computational complexity, and constraining the boundaries of the heart image through dual-task learning.
The performance of full-heart segmentation is improved, the number of parameters and calculations of the model is reduced, and the segmentation effect is significantly improved through boundary constraints, which is better than other comparison methods.
Smart Images

Figure CN117011306B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of medical image processing and computer vision. Background Art
[0002] Heart segmentation is the basis and key to the analysis and treatment of cardiovascular diseases, and plays an important role in the medical field. Extracting anatomical structures from heart images can assist doctors in planning surgeries in advance and ensure the smooth implementation of surgeries. At present, accurate and efficient whole-heart segmentation remains a challenging task. Although fully convolutional neural networks (FCNNs) have made remarkable achievements in the field of medical segmentation, a number of excellent algorithms have emerged. However, due to the limitation of the receptive field, FCNNs often find it difficult to capture the global information of the image, resulting in insufficient global spatial feature modeling, and the segmentation performance needs to be further improved.
[0003] Given the outstanding performance of Transformer in natural language processing, it has been widely used in visual tasks and achieved remarkable results, such as image recognition, object detection, image generation, etc. As a sequence-to-sequence model based on the self-attention mechanism, Transformer can well model the relationship between image blocks and obtain the global features of the image. However, the self-attention mechanism usually has a large computational complexity, and training a 3D Transformer model is very resource-intensive. At present, many studies are committed to exploring the self-attention mechanism of Transformer. For example, Swin Transformer uses a translation window operation to confine the calculation of the self-attention mechanism to the window, and realizes the interaction of global features through a cross-window method. This self-attention calculation method greatly reduces the computational complexity and has significant effects on various segmentation tasks, but in 3D Transformer, it is a great challenge to resources. Therefore, seeking a simple self-attention mechanism for the whole heart segmentation task is the focus of our exploration.
[0004] In addition, since cardiac images are affected by noise, low contrast and other factors, the boundaries of cardiac structures are blurred and difficult to distinguish. Therefore, it is necessary to impose additional constraints on the boundaries of the heart. It can enhance the model's ability to learn boundaries and further improve segmentation performance. Summary of the invention
[0005] The present invention aims to solve one of the technical problems in the related art at least to a certain extent.
[0006] To this end, the purpose of the present invention is to propose a multi-task whole heart CT segmentation method based on pooling Transformer for accurate and efficient heart segmentation.
[0007] To achieve the above objectives, the first embodiment of the present invention proposes a multi-task whole heart CT segmentation method based on pooling Transformer, comprising:
[0008] Obtain a whole heart CT image training dataset;
[0009] Constructing a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and uses a 3D pooling attention mechanism to calculate the relationship between image blocks;
[0010] Inputting the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; wherein the segmentation loss function is composed of a Dice loss function and a weighted cross entropy loss function;
[0011] The images in the test data set are input into the trained whole heart CT image segmentation model to obtain the image segmentation results.
[0012] In addition, the multi-task whole heart CT segmentation method based on pooling Transformer according to the above embodiment of the present invention may also have the following additional technical features:
[0013] Furthermore, in one embodiment of the present invention, the 3D transformer encoder and the 3D convolutional neural network decoder are connected in a skip-layer manner.
[0014] Furthermore, in one embodiment of the present invention, before the whole heart CT image training data set is input into the whole heart CT image segmentation model, the model is converted into an embedded representation, specifically including:
[0015] Let the input image x∈R W×H×D×C , respectively represent the width, height, number of slices and number of channels of the slice, the number of channels is 1; the input image is divided into non-overlapping image blocks of the same size, the i-th image block The size is P×P×P, the number of channels is C, and all image blocks consist of X={x1,x2,…,x N ,}, N represents the number of image blocks;
[0016] X is mapped to a K-dimensional embedding representation through a linear layer, the positional encoding is added to the input vector, and input into the Transformer. This process is formalized as follows:
[0017]
[0018] Among them, Epos represents the position code, Represents a spatial mapping.
[0019] Image features are extracted through Transformer, and the downsampling operation reduces the resolution to half of the original resolution to form a multi-scale feature table, where the operation of each Transformer block is expressed as follows:
[0020] z′ l =MSA(Norm(z i-1 ))+z i-1 ,
[0021] z l =MLP(Norm(z′ l ))+z′ l ,
[0022] Among them, MSA(·) represents multi-head self-attention, Norm(·) represents layer normalization, MLP(·) represents multi-layer perceptron, z i-1 represents the output of the previous layer; specifically, the output information of the previous layer is first normalized to a data distribution with a mean of 0 and a variance of 1, and then the relationship between the image blocks is calculated using a multi-head self-attention mechanism, and the obtained output is connected to z by a residual connection. i-1 Add to get z′ i .
[0023] Furthermore, in one embodiment of the present invention, the method of calculating the relationship between image blocks using a multi-head self-attention mechanism includes:
[0024] Combine the input features with the weight matrix W Q , W K , W V Multiply them to get the projection of Q, K, and V, dot-multiply Q and K to get the attention score between the image blocks, divide the attention score by a constant, and normalize the score to measure the correlation between the two image blocks, which is expressed as follows:
[0025]
[0026] The attention score is used as a weight coefficient and multiplied by V to get the output, which can be expressed as:
[0027] SA(W,V)=WV,
[0028] Multiple single-head self-attention mechanisms are spliced together to form a multi-head self-attention mechanism:
[0029] MA(Q,K,V)=Concat(h1,h2,…,hh )W o
[0030] h i =SA(W i ,V i );
[0031] The multi-head self-attention mechanism is simplified, and the 3D pooling operation is used to replace the above complex multi-head self-attention calculation.
[0032] Furthermore, in one embodiment of the present invention, an auxiliary task is added to the whole heart CT image segmentation model to constrain the boundary, and the conversion from the real mask label to the directional distance map is realized by the following formula:
[0033]
[0034] Where S in , S out They represent the inner area of the heart, the boundary of the heart and the outer area of the heart respectively.
[0035] Furthermore, in one embodiment of the present invention, the segmentation loss function of the whole heart CT image segmentation model adopts the Dice loss function L dice And the weighted cross entropy loss function L wce Composition. They are expressed by calculation formulas as follows:
[0036]
[0037]
[0038] L seg =L dice +L wce ,
[0039] in, represents the true label, represents the predicted value of the i-th voxel predicted as category c, N represents the number of voxels, and c represents the number of categories;
[0040] The boundary regression loss function of the whole heart CT image segmentation model adopts mean square error, which is expressed as follows:
[0041]
[0042] Among them, d i represents the SDM label converted from the segmentation truth value, is the predicted value output by the model;
[0043] The final loss function L totalIt is composed of the weighted segmentation loss and boundary regression loss function, with a weighting coefficient of 1.5, as shown below:
[0044] L total =L seg +1.5L sdm .
[0045] To achieve the above objectives, the second embodiment of the present invention proposes a multi-task whole heart CT segmentation device based on pooling Transformer, comprising the following modules:
[0046] An acquisition module, used for acquiring a whole heart CT image training data set;
[0047] A construction module is used to construct a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and a 3D pooling attention mechanism is used to calculate the relationship between image blocks;
[0048] A training module, used for inputting the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; wherein the segmentation loss function is composed of a Dice loss function and a weighted cross entropy loss function;
[0049] The output module is used to input the image to be segmented into the trained whole heart CT image segmentation model to obtain the image segmentation result.
[0050] To achieve the above-mentioned purpose, the third aspect of the present invention proposes a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the multi-task whole heart CT segmentation method based on pooling Transformer as described above.
[0051] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes a computer-readable storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, a multi-task whole heart CT segmentation method based on pooling Transformer as described above is implemented.
[0052] The multi-task whole heart CT segmentation method based on pooling transformer proposed in the embodiment of the present invention improves the segmentation performance while reducing the number of model parameters and calculations; at the same time, a boundary-constrained dual-task learning method is proposed, with the segmentation task as the main task and the boundary task as the auxiliary task, which constrains the boundaries of the heart image and further improves the segmentation performance. Experiments conducted on the MM-WHS whole heart CT dataset show that the method proposed in the present invention is superior to other comparative methods, proving the effectiveness of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0054] Figure 1 A flowchart of a multi-task whole heart CT segmentation method based on pooling Transformer provided in an embodiment of the present invention.
[0055] Figure 2 A schematic diagram of the overall framework of a model provided in an embodiment of the present invention.
[0056] Figure 3 A schematic diagram of a multi-head self-attention mechanism provided in an embodiment of the present invention.
[0057] Figure 4 A schematic diagram of a 3D average pooling operation provided by an embodiment of the present invention.
[0058] Figure 5 A schematic flow chart of a multi-task whole heart CT segmentation device based on a pooling Transformer provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0059] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0060] The following describes a multi-task whole heart CT segmentation method based on pooling Transformer according to an embodiment of the present invention with reference to the accompanying drawings.
[0061] Figure 1 A flowchart of a multi-task whole heart CT segmentation method based on pooling Transformer provided in an embodiment of the present invention.
[0062] like Figure 1As shown in FIG. 1 , the multi-task whole heart CT segmentation method based on pooling Transformer includes the following steps:
[0063] S101: Acquire a whole heart CT image training dataset;
[0064] S102: constructing a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and uses a 3D pooling attention mechanism to calculate the relationship between image blocks;
[0065] S103: inputting the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; wherein the segmentation loss function is composed of a Dice loss function and a weighted cross entropy loss function;
[0066] S104: Input the images in the test data set into the trained whole heart CT image segmentation model to obtain image segmentation results.
[0067] The overall framework of the model is as follows Figure 2 As shown in the figure, it is mainly divided into two parts: 3D transformer encoder and 3D convolutional neural network decoder.
[0068] Furthermore, in one embodiment of the present invention, the 3D transformer encoder and the 3D convolutional neural network decoder adopt a skip-layer connection method.
[0069] Furthermore, in one embodiment of the present invention, before the whole heart CT image training data set is input into the whole heart CT image segmentation model, the model is converted into an embedded representation, specifically including:
[0070] Let the input image x∈R W×H×D×C , respectively represent the width, height, number of slices and number of channels of the slice, the number of channels is 1;
[0071] The input image is divided into non-overlapping image blocks of the same size. The i-th image block The size is P×P×P, the number of channels is C, and all image blocks consist of X={x1,x2,…,x N ,}, N represents the number of image blocks;
[0072] X is mapped to a K-dimensional embedding representation through a linear layer, the positional encoding is added to the input vector, and input into the Transformer. This process is formalized as follows:
[0073]
[0074] Among them, E pos represents the position code, Represents a spatial mapping.
[0075] Image features are extracted through Transformer, and the downsampling operation reduces the resolution to half of the original resolution to form a multi-scale feature table, where the operation of each Transformer block is expressed as follows:
[0076] z′ l =MSA(Norm(z i-1 ))+z i-1 ,
[0077] z l =MLP(Norm(z′ l ))+z′ l ,
[0078] Among them, MSA(·) represents multi-head self-attention, Norm(·) represents layer normalization, MLP(·) represents multi-layer perceptron, z i-1 represents the output of the previous layer; specifically, the output information of the previous layer is first normalized to a data distribution with a mean of 0 and a variance of 1, and then the relationship between the image blocks is calculated using a multi-head self-attention mechanism, and the obtained output is connected to z by a residual connection. i-1 Add to get z′ i .
[0079] Specifically, the 3D transformer encoder is composed of multiple stacked Transformer blocks. In this work, a total of 8 Transformer blocks are used, divided into four stages. Each stage contains two Transformer blocks and a downsampling operation. The Transformer block is used to extract image features. The downsampling operation reduces the resolution to half of the original to form a multi-scale feature representation. Each Transformer block consists of layer normalization, multi-head self-attention mechanism, and multi-layer perceptron. In order to further abstract the features and improve the representation ability of the features, a multi-layer perceptron is used, that is, two linear layers are further abstracted and strengthened, and the previous features are retained again through residual connections.
[0080] For the 8 stacked encoders, the outputs of the 2nd, 4th, 6th and 8th layers are obtained respectively. A 3D convolutional neural network is connected after the output of the four layers to map the embedding space to the input space for further connection with the decoder of the same scale. The decoder is symmetrical with the encoder and consists of multi-layer convolution blocks and upsampling. The upsampling adopts the deconvolution method. Concat is used for skip-layer connection between the encoder and the decoder.
[0081] The core of Transformer is the self-attention mechanism, which is used to calculate the relationship between image blocks and model their global information. The traditional Transformer contains multiple parallel single-head self-attention. Figure 3 shown.
[0082] Furthermore, in one embodiment of the present invention, a multi-head self-attention mechanism is used to calculate the relationship between image blocks, including:
[0083] Combine the input features with the weight matrix W Q , W K , W V Multiply them to get the projection of Q, K, and V. Dot-multiply Q and K to get the attention score between the image blocks. Divide the attention score by a constant and normalize the score to measure the correlation between the two image blocks. The calculation formula is:
[0084]
[0085] The attention score is used as a weight coefficient and multiplied by V to get the output, which can be expressed as:
[0086] SA(W,V)=WV,
[0087] Multiple single-head self-attention mechanisms are spliced together to form a multi-head self-attention mechanism:
[0088] MA(Q,K,V)=Concat(h1,h2,…,h h )W o
[0089] h i =SA(W i ,V i );
[0090] The multi-head self-attention mechanism is simplified, and the 3D pooling operation is used to replace the above complex multi-head self-attention calculation.
[0091] As the core component of Transformer, the multi-head self-attention mechanism has a quadratic computational complexity and consumes a lot of resources when processing 3D images. The multi-head self-attention mechanism is essentially a module that processes the relationship between image blocks. To improve the problem of complex computation, we extend the 2D pooling operation to the 3D pooling operation and use 3D pooling as a self-attention mechanism in the heart segmentation task to simply fuse the relationship between image blocks. Specifically, we explored the role of the 3D average pooling self-attention mechanism and the 3D maximum self-attention mechanism in the heart segmentation task. Taking average pooling as an example, Figure 4As shown. Since the calculation of the self-attention relationship in the Transformer does not involve changes in feature resolution, an operation with a padding of 1, a window size of 3, and a step size of 1 is used. The average of the voxels in the window is extracted each time as the output, fully considering the voxels inside the window and retaining the information of all voxels to the greatest extent. Considering that a large number of voxels are redundant, the problem of unclear extracted features may occur. The self-attention mechanism is further improved to a maximum pooling operation, which outputs the maximum value in the window and retains the most significant information. In short, 3D pooling can, in principle, fuse the features of each window and extract the relationship between each image block. As a self-attention mechanism, it can minimize the amount of calculation and parameters.
[0092] The grayscale values of the cardiac image are similar at the border, showing a fuzzy boundary phenomenon, which makes it very difficult for the model to classify the boundary area. To address this problem, the present invention adopts a dual-task learning approach, adding an auxiliary task in addition to the segmentation task to constrain the boundary.
[0093] Furthermore, in one embodiment of the present invention, an auxiliary task is added to the whole heart CT image segmentation model to constrain the boundary, and the conversion from the true mask label to the directional distance map is realized by the following formula:
[0094]
[0095] Where S in , S out They represent the inner area of the heart, the boundary of the heart and the outer area of the heart respectively.
[0096] The value of the voxel is determined by the distance from the boundary. The voxel value at the boundary is 0, the inside of the heart is represented by a negative value, and the outside is represented by a positive value. The farther from the boundary, the smaller the absolute value of the voxel.
[0097] Furthermore, in one embodiment of the present invention, the segmentation loss function of the whole heart CT image segmentation model adopts the Dice loss function L dice And the weighted cross entropy loss function L wce Composition. They are expressed by calculation formulas as follows:
[0098]
[0099]
[0100] L seg =L dice +L wce ,
[0101] in, represents the true label, represents the predicted value of the i-th voxel predicted as category c, N represents the number of voxels, and c represents the number of categories;
[0102] The boundary regression loss function of the whole heart CT image segmentation model uses the mean square error, which is expressed as follows:
[0103]
[0104] Among them, d i represents the SDM label converted from the segmentation truth value, is the predicted value output by the model;
[0105] The final loss function L total It is composed of the weighted segmentation loss and boundary regression loss function, with a weighting coefficient of 1.5, as shown below:
[0106] L total =L seg +1.5L sdm .
[0107] Among them, the weight coefficient aims to balance the weight of the number of categories, giving more weight to categories with fewer voxels to better handle the imbalance of samples.
[0108] The embodiment of the present invention proposes a multi-task whole heart CT segmentation method based on pooling Transformer, with 3D Transformer as the encoder and CNN as the decoder. 3D average pooling and 3D maximum pooling are used as the self-attention mechanism of Transformer, which not only reduces the number of parameters and calculations, but also improves the segmentation performance of the model. At the same time, in order to further improve the problem of difficult boundary classification, a dual-task learning method is proposed, which uses the boundary regression task as an auxiliary task and the segmentation task as the main task to strengthen the constraint on the boundary to further improve the segmentation effect. The proposed method was experimented on the whole heart CT dataset, and the results showed that it was better than other methods, proving the effectiveness of the method.
[0109] In order to implement the above embodiment, the present invention also proposes a multi-task whole heart CT segmentation device based on pooling Transformer.
[0110] Figure 5 A schematic diagram of the structure of a multi-task whole heart CT segmentation device based on a pooling Transformer provided in an embodiment of the present invention.
[0111] like Figure 5 As shown, the multi-task whole heart CT segmentation device based on pooling Transformer includes: an acquisition module 100, a construction module 200, a training module 300, and an output module 400, wherein:
[0112] An acquisition module, used for acquiring a whole heart CT image training data set;
[0113] A construction module is used to construct a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and uses a 3D pooling attention mechanism to calculate the relationship between image blocks;
[0114] A training module is used to input the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; wherein the segmentation loss function is composed of a Dice loss function and a weighted cross entropy loss function;
[0115] The output module is used to input the image to be segmented into the trained whole heart CT image segmentation model to obtain the image segmentation result.
[0116] To achieve the above-mentioned purpose, the third aspect of the present invention proposes a computer device, characterized in that it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, it implements the multi-task whole heart CT segmentation method based on pooling Transformer as described above.
[0117] To achieve the above-mentioned purpose, the fourth aspect of the present invention proposes a computer-readable storage medium on which a computer program is stored, characterized in that when the computer program is executed by a processor, the multi-task whole heart CT segmentation method based on pooling Transformer as described above is implemented.
[0118] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0119] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present invention, the meaning of "plurality" is at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0120] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A multi-task whole heart CT segmentation method based on pooling Transformer, characterized in that: The following steps are involved: Obtain a whole heart CT image training dataset; Constructing a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and uses a 3D pooling attention mechanism to calculate the relationship between image blocks; Inputting the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; The segmentation loss function consists of the Dice loss function and the weighted cross entropy loss function; Input the images in the test data set into the trained whole heart CT image segmentation model to obtain the image segmentation results; The 3D pooling attention mechanism is used to calculate the relationship between image blocks, including: Combine the input features with the weight matrix , , Multiply them to get the projection of Q, K, and V, dot-multiply Q and K to get the attention score between the image blocks, divide the attention score by a constant, and normalize the score to measure the correlation between the two image blocks, which is expressed as follows: , The attention score is used as a weight coefficient and multiplied by V to get the output, which can be expressed as: , Multiple single-head self-attention mechanisms are spliced together to form a multi-head self-attention mechanism: ; Simplify the multi-head self-attention mechanism and use 3D pooling operation to replace the complex multi-head self-attention calculation. The segmentation loss function of the whole heart CT image segmentation model adopts the Dice loss function And the weighted cross entropy loss function The composition is expressed by the following calculation formulas: , , , in, represents the true label, represents the predicted value of the i-th voxel predicted as category c, N represents the number of voxels, and c represents the number of categories; The boundary regression loss function of the whole heart CT image segmentation model adopts mean square error, which is expressed as follows: , in, represents the SDM label converted from the segmentation truth value, is the predicted value output by the model; Final loss function It is composed of the weighted segmentation loss and boundary regression loss function, with a weighting coefficient of 1.5, as shown below: 。 2. The method according to claim 1, characterized in that: The 3D transformer encoder and the 3D convolutional neural network decoder are connected in a skip-layer manner.
3. The method according to claim 1, characterized in that Before inputting the whole heart CT image training data set into the whole heart CT image segmentation model, converting the model into an embedded representation, specifically including: Assume the input image , respectively represent the width, height, number of slices and number of channels of the slice, the number of channels is 1; The input image is divided into non-overlapping image blocks of the same size, the i-th image block , indicating the size is , the number of channels is , all image blocks consist of A sequence of, N represents the number of image blocks; Will Through the linear layer mapping The positional encoding is added to the input vector and fed into the Transformer. This process is formalized as follows: , in, represents the position code, Represents spatial mapping; Image features are extracted through Transformer, and the downsampling operation reduces the resolution to half of the original resolution to form a multi-scale feature table, where the operation of each Transformer block is expressed as follows: , , in, represents multi-head self-attention, Representation layer normalization, represents a multi-layer perceptron, represents the output of the previous layer; specifically, the output information of the previous layer is first normalized to a data distribution with a mean of 0 and a variance of 1, and then the relationship between the image blocks is calculated using a multi-head self-attention mechanism, and the obtained output is connected to Add together to get .
4. The method according to claim 1, characterized in that: The method also includes adding an auxiliary task to the whole heart CT image segmentation model to constrain the boundary, and realizing the conversion from the real mask label to the directional distance map through the following formula: , in , , They represent the inner area of the heart, the boundary of the heart and the outer area of the heart respectively.
5. A multi-task whole heart CT segmentation device based on pooling Transformer, characterized in that: Includes the following modules: An acquisition module, used for acquiring a whole heart CT image training data set; A construction module is used to construct a whole heart CT image segmentation model, wherein the whole heart CT image segmentation model includes a 3D transformer encoder and a 3D convolutional neural network decoder, and uses a 3D pooling attention mechanism to calculate the relationship between image blocks; A training module, used for inputting the whole heart CT image training data set into the whole heart CT image segmentation model for training to obtain a completed whole heart CT image segmentation model; The segmentation loss function consists of the Dice loss function and the weighted cross entropy loss function; An output module is used to input the image to be segmented into the trained whole heart CT image segmentation model to obtain the image segmentation result; The 3D pooling attention mechanism is used to calculate the relationship between image blocks, including: Combine the input features with the weight matrix , , Multiply them to get the projection of Q, K, and V, dot-multiply Q and K to get the attention score between the image blocks, divide the attention score by a constant, and normalize the score to measure the correlation between the two image blocks, which is expressed as follows: , The attention score is used as a weight coefficient and multiplied by V to get the output, which can be expressed as: , Multiple single-head self-attention mechanisms are spliced together to form a multi-head self-attention mechanism: ; Simplify the multi-head self-attention mechanism and use 3D pooling operation to replace the complex multi-head self-attention calculation. The segmentation loss function of the whole heart CT image segmentation model adopts the Dice loss function And the weighted cross entropy loss function The composition is expressed by the following calculation formulas: , , , in, represents the true label, represents the predicted value of the i-th voxel predicted as category c, N represents the number of voxels, and c represents the number of categories; The boundary regression loss function of the whole heart CT image segmentation model adopts mean square error, which is expressed as follows: , in, represents the SDM label converted from the segmentation truth value, is the predicted value output by the model; Final loss function It is composed of the weighted segmentation loss and boundary regression loss function, with a weighting coefficient of 1.5, as shown below: 。 6. A computer device, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-task whole heart CT segmentation method based on pooling Transformer as described in any one of claims 1 to 4.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the multi-task whole heart CT segmentation method based on pooling Transformer as described in any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Auxiliary detection method and image recognition method for rib fractures based on deep learning
US20220198230A1
Image segmentation method, apparatus and device, and storage medium
WO2022032823A1