Breast tumor accurate segmentation method based on space-time fusion network

The spatiotemporal fusion network addresses low precision and interpretability issues in breast tumor segmentation by integrating residual convolution, LSTM, and PK feature fusion, resulting in high-precision and medically interpretable tumor segmentation.

CN120318247APending Publication Date: 2025-07-15GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510382937.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing deep learning methods have low segmentation accuracy and lack medical explanatory in breast cancer DCE-MRI image segmentation.

Method used

A space-time fusion network based on Unet structure is adopted, combining breast DCE-MRI data and PK feature maps, and spatial and temporal features are captured through residual convolution, LSTM and SE channel attention mechanisms, improving segmentation accuracy and enhancing medical interpretability.

Benefits of technology

It realizes high-precision segmentation of breast tumors and provides medically interpretable tumor characteristics to support doctors' more accurate diagnosis and treatment planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318247A_ABST
    Figure CN120318247A_ABST
Patent Text Reader

Abstract

The invention discloses a breast tumor accurate segmentation method based on a space-time fusion network, which adopts the space-time fusion network fusing convolution, LSTM and PK feature fusion to capture space and time features at the same time so as to improve segmentation precision and enhance medical interpretability, and introduces LSTM into each layer of an encoder of the space-time fusion network to capture sequence dependence, so that the segmentation precision is improved. And the model can learn the mode in the whole dynamic enhancement process. In addition, parameters of the classic PK model are fused into network input and training processes, so that the medical interpretability of the model is improved. By comprehensively utilizing a deep learning framework of DCE-MRI spatio-temporal dynamic information and radiomics characteristics, high-precision segmentation of breast tumors is realized, and medically interpretable tumor characteristics are provided, so that doctors can make more accurate diagnosis and treatment plans based on segmentation results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of medical image processing and artificial intelligence, and particularly relates to a precise segmentation method for breast tumors based on a spatio-temporal fusion network. Background Art

[0002] Breast cancer is one of the most common malignant tumors among women globally, and its early detection and precise diagnosis are crucial for improving the prognosis of patients. Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) has been widely used in breast cancer detection because it can provide hemodynamic information of tumors. However, due to the complex morphology of tumors, diverse enhancement patterns, and noise interference during the imaging process, traditional image segmentation methods are difficult to meet clinical needs.

[0003] With the development of deep learning, deep learning methods have demonstrated powerful capabilities in the field of medical image segmentation. Currently, most traditional deep learning segmentation methods utilize single-phase images or simply compare pre- and post-contrast images, without fully exploiting the temporal dimension features. Therefore, it is often difficult to improve the segmentation accuracy. Moreover, traditional deep learning segmentation methods are often regarded as "black boxes" and lack medical interpretability. Summary of the Invention

[0004] The problem to be solved by the present invention is that existing deep learning methods have poor segmentation accuracy and medical interpretability. The present invention provides a precise segmentation method for breast tumors based on a spatio-temporal fusion network.

[0005] To solve the above problems, the present invention is realized through the following technical solutions:

[0006] A precise segmentation method for breast tumors based on a spatio-temporal fusion network includes the following steps:

[0007] Step 1: Construct a spatio-temporal fusion network based on the Unet structure;

[0008] The spatio-temporal fusion network consists of an input layer, 4 encoder units, 3 downsampling units, 3 attention skip connection units, 4 upsampling units, 4 decoder units, and an output layer; the input of the input layer forms the main path input of the spatio-temporal fusion network, and the auxiliary inputs of the first encoder unit, the second encoder unit, the third encoder unit, and the fourth encoder unit form the auxiliary input of the spatio-temporal fusion network; the output of the input layer is connected to the main path input of the first encoder unit, the output of the first encoder unit is connected to the main path input of the second encoder unit through the first downsampling unit, the output of the second encoder unit is connected to the main path input of the third encoder unit through the second downsampling unit, and the output of the third encoder unit is connected to the main path input of the fourth encoder unit through the third downsampling unit; the output of the fourth encoder unit is connected to the main path input of the first decoder unit through the first upsampling unit, the output of the first decoder unit is connected to the main path input of the second decoder unit through the second upsampling unit, the output of the second decoder unit is connected to the main path input of the third decoder unit through the third upsampling unit, and the output of the third decoder unit is connected to the main path input of the fourth decoder unit through the fourth upsampling unit; the output of the first encoder unit is also connected to the skip input of the third decoder unit through the first attention skip connection unit, the output of the second encoder unit is also connected to the skip input of the second decoder unit through the second attention skip connection unit, the output of the third encoder unit is also connected to the skip input of the first decoder unit through the third attention skip connection unit, and the output of the fourth decoder unit is connected to the input of the output layer; the output of the output layer forms the output of the spatio-temporal fusion network;

[0009] Step 2: Obtain the breast DCE-MRI dataset, and use the extended Tofts model to extract the breast PK feature maps corresponding to each group of breast DCE-MRI temporal image maps in the breast DCE-MRI dataset to obtain the breast PK feature map dataset;

[0010] Step 3: Use the breast DCE-MRI dataset and the breast PK feature map dataset to train and validate the spatio-temporal fusion network to obtain an accurate breast tumor segmentation model;

[0011] Step 4: Obtain the breast DCE-MRI to be segmented, and use the extended Tofts model to extract the breast PK feature map to be segmented corresponding to the breast DCE-MRI to be segmented;

[0012] Step 5: Send the breast DCE-MRI to be segmented and the breast PK feature map to be segmented into the accurate breast tumor segmentation model to obtain the segmentation result.

[0013] In the above solution, both the first encoder unit and the fourth encoder unit of the spatio-temporal fusion network are composed of 3 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 3 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the third residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit. The second encoder unit of the spatio-temporal fusion network is composed of 4 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 4 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the fourth residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit. The third encoder unit of the spatio-temporal fusion network is composed of 6 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 6 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the sixth residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit.

[0014] In the above solution, each decoder unit of the spatio-temporal fusion network is composed of a fusion layer, a convolutional layer, and a double-layer residual convolutional layer; one input of the fusion layer forms the main path input of this decoder unit. The other input of the fusion layer forms the skip input of this decoder unit. The output of the fusion layer is connected to the input of the convolutional layer. The output of the convolutional layer is connected to the input of the double-layer residual convolutional layer. The output of the double-layer residual convolutional layer forms the output of this decoder unit.

[0015] In the above solution, the attention skip connection unit of the spatio-temporal fusion network is a SE channel attention unit.

[0016] In the above solution, the input layer of the spatio-temporal fusion network is composed of a 7×7 convolutional layer and a 3×3 pooling layer; the input of the 7×7 convolutional layer forms the input of the input layer. The output of the 7×7 convolutional layer is connected to the input of the 3×3 pooling layer. The output of the 3×3 pooling layer forms the output of the input layer.

[0017] In the above solution, the output layer of the spatio-temporal fusion network is composed of a bilinear interpolation upsampling layer, a 1×1 convolutional layer, and a Sigmoid activation function layer; the input of the bilinear interpolation upsampling layer forms the input of the output layer. The output of the bilinear interpolation upsampling layer is connected to the input of the 1×1 convolutional layer. The output of the 1×1 convolutional layer is connected to the input of the Sigmoid activation function layer. The output of the Sigmoid activation function layer forms the output of the output layer.

[0018] In the above step 3, the breast DCE-MRI dataset is input through the main path of the spatio-temporal fusion network, and the breast PK feature map dataset is input through the auxiliary input of the spatio-temporal fusion network.

[0019] In the above step 5, the breast DCE-MRI to be segmented is input through the main path of the breast tumor precise segmentation model, and the breast PK feature map to be segmented is input through the auxiliary input of the breast tumor precise segmentation model.

[0020] Compared with the prior art, the present invention adopts a spatio-temporal fusion network that integrates convolution, LSTM, and PK feature fusion to simultaneously capture spatial and temporal features, aiming to improve the segmentation accuracy and enhance medical interpretability. Considering that breast tumor DCE-MRI provides 4D information on tumor enhancement over time, but many existing methods only utilize single-phase images or simply adopt pre- and post-contrast images, without fully exploiting the features in the time dimension, the present invention introduces recurrent units (LSTM) in each layer of the encoder to capture sequence dependencies, enabling the model to learn the patterns in the entire dynamic enhancement process. In addition, considering that traditional deep learning segmentation is often regarded as a "black box" and lacks medical interpretability, the present invention incorporates the parameters of the classical PK model into the network input and training process to improve the medical interpretability of the model. The present invention realizes high-precision segmentation of breast tumors through a deep learning framework that comprehensively utilizes the spatio-temporal dynamic information of DCE-MRI and radiomics features, and provides medically interpretable tumor features, enabling doctors to make more accurate diagnoses and treatment plans based on the segmentation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic diagram of the principle of the spatio-temporal fusion network based on the Unet structure.

[0022] Figure 2 It is a schematic diagram of the principle of the encoder unit.

[0023] Figure 3 It is a schematic diagram of the principle of the decoder unit.

[0024] Figure 4 It is a 3D time series diagram matrix of DEC-MRI.

[0025] Figure 5 It is a schematic diagram for extracting the time-signal intensity curve.

[0026] Figure 6 It is a schematic diagram for fitting the PK feature map. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to specific examples and the accompanying drawings.

[0028] A precise segmentation method for breast tumors based on a spatio-temporal fusion network, which comprises the following steps:

[0029] Step 1, construct a spatio-temporal fusion network based on the Unet structure.

[0030] See Figure 1 , the constructed spatio-temporal fusion network is composed of an input layer, 4 encoder units, 3 downsampling units, 3 attention skip connection units, 4 upsampling units, 4 decoder units, and an output layer. The input of the input layer forms the main path input of the spatio-temporal fusion network, and the auxiliary inputs of the first encoder unit, the second encoder unit, the third encoder unit, and the fourth encoder unit form the auxiliary input of the spatio-temporal fusion network. The output of the input layer is connected to the main path input of the first encoder unit, the output of the first encoder unit is connected to the main path input of the second encoder unit through the first downsampling unit, the output of the second encoder unit is connected to the main path input of the third encoder unit through the second downsampling unit, and the output of the third encoder unit is connected to the main path input of the fourth encoder unit through the third downsampling unit. The output of the fourth encoder unit is connected to the main path input of the first decoder unit through the first upsampling unit, the output of the first decoder unit is connected to the main path input of the second decoder unit through the second upsampling unit, the output of the second decoder unit is connected to the main path input of the third decoder unit through the third upsampling unit, and the output of the third decoder unit is connected to the main path input of the fourth decoder unit through the fourth upsampling unit. The output of the first encoder unit is also connected to the skip input of the third decoder unit through the first attention skip connection unit, the output of the second encoder unit is also connected to the skip input of the second decoder unit through the second attention skip connection unit, the output of the third encoder unit is also connected to the skip input of the first decoder unit through the third attention skip connection unit, and the output of the fourth decoder unit is connected to the input of the output layer. The output of the output layer forms the output of the spatio-temporal fusion network.

[0031] (1) Input layer

[0032] The input layer is composed of a 7×7 convolutional layer and a 3×3 pooling layer. The input of the 7×7 convolutional layer forms the input of the input layer, the output of the 7×7 convolutional layer is connected to the input of the 3×3 pooling layer, and the output of the 3×3 pooling layer forms the output of the input layer.

[0033] (2) Encoder

[0034] The encoder, as the first half of the entire network, is responsible for extracting spatial and temporal information in multi-phase DCE-MRI sequences. Its structure is designed with a triple mechanism of residual convolution-LSTM temporal modeling-PK fusion. The encoder uses residual convolution to extract spatial features, LSTM to capture temporal dependencies, and fuses PK feature maps to enhance the interpretability of encoding.

[0035] The residual convolution unit uses ResNet-34 as the backbone network. The global average pooling layer and fully connected layer of ResNet-34 are removed, and only the convolutional layer is retained. It contains five stages: The first stage is the initial convolutional layer: a 7×7 convolutional kernel, a stride of 2, 64 output channels, followed by batch normalization (BatchNorm) and ReLU activation function, and then a 3×3 max pooling layer (stride 2); Each subsequent stage is as follows: In the second stage, 3 residual blocks are used, the pooling layer has a stride of 1, and the output channels are 64; In the third stage, 4 residual blocks, the pooling layer has a stride of 2, and the output channels are 128; In the fourth stage, 6 residual blocks, the pooling layer has a stride of 2, and the output channels are 256; In the fifth stage, 3 residual blocks, the pooling layer has a stride of 2, and the output channels are 512; The size of the output feature map in each stage is reduced to 1 / 4, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original image size in turn.

[0036] LSTM (Long Short-Term Memory) is embedded at the end of the residual convolution unit to process the multi-temporal feature sequence. The feature map (size B×T×C×H×W) output in the current stage is unfolded into a sequence by time steps and input into an LSTM with a hidden layer dimension of C, and the temporally enhanced feature with a size of B×T×C×H×W is output. Among them, B is the training batch size, T is the number of time phases, C is the number of feature channels, H is the height of the feature map, and W is the height of the feature map.

[0037] The fusion layer will copy the PK feature map with a size of B×3×H×W to each time step to make its size become B×T×3×H×W, and then concatenate it with the temporally enhanced feature output by the LSTM along the channel dimension to form a feature tensor with a size of B×T×(C + 3)×H×W. A 3D convolution operation (kernel size 3×3×3) is used to convolve and integrate this five-dimensional feature tensor to fuse the spatial, temporal, and PK parameter information, and the output size is B×T×C×H×W. Among them, B is the training batch size, T is the number of time phases, C is the number of feature channels, H is the height of the feature map, and W is the height of the feature map.

[0038] In the present invention, the 4 encoder units adopt similar structures, such as Figure 2As shown. The structures of the first encoder unit and the fourth encoder unit are the same, both being three-residual encoder units, that is, both consisting of 3 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 3 residual convolutional layers are connected in series in sequence, the input of the first residual convolutional layer forms the main path input of this encoder unit, the output of the third residual convolutional layer is connected to the input of the LSTM, the output of the LSTM is connected to one input of the fusion layer, the other input of the fusion layer forms the auxiliary input of this encoder unit, and the output of the fusion layer forms the output of this encoder unit. The second encoder unit is a four-residual encoder unit, that is, consisting of 4 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 4 residual convolutional layers are connected in series in sequence, the input of the first residual convolutional layer forms the main path input of this encoder unit, the output of the fourth residual convolutional layer is connected to the input of the LSTM, the output of the LSTM is connected to one input of the fusion layer, the other input of the fusion layer forms the auxiliary input of this encoder unit, and the output of the fusion layer forms the output of this encoder unit. The third encoder unit is a six-residual encoder unit, that is, consisting of 6 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 6 residual convolutional layers are connected in series in sequence, the input of the first residual convolutional layer forms the main path input of this encoder unit, the output of the sixth residual convolutional layer is connected to the input of the LSTM, the output of the LSTM is connected to one input of the fusion layer, the other input of the fusion layer forms the auxiliary input of this encoder unit, and the output of the fusion layer forms the output of this encoder unit.

[0039] (3) Attention skip connection

[0040] To further improve the effectiveness of the features transmitted in the skip connection, the present invention introduces a channel attention mechanism (Channel Attention) in the skip connection of the UNet structure to enhance the key feature channels and suppress the redundant channels, thereby improving the discriminative ability of the tumor region and the boundary sensitivity. In the present invention, the attention skip connection unit is an SE channel attention unit. For the feature map F e output by the encoder in each pair of skip connections, an SE channel attention unit is introduced to perform weighted adjustment on it to obtain the enhanced encoder feature map e ′ F . Subsequently, the enhanced feature map e ′ F is concatenated with the feature map F d obtained by upsampling the decoder in the channel dimension: This feature map will be used for the recovery of high-resolution features. By introducing the channel attention mechanism, the network can effectively highlight the tumor-related feature channels and weaken the irrelevant background information, thereby improving the segmentation accuracy, edge continuity, and the model's perception ability for small targets.

[0041] (4) Decoder

[0042] As the latter half of the entire network, the decoder's function is to gradually restore the spatial resolution and enhance the expression ability of the tumor region through step-by-step upsampling and feature enhancement mechanisms. Its structure is designed as a skip connection fusion - double-layer residual convolution dual mechanism. The decoder fuses skip connection features and extracts local spatial context information.

[0043] The fusion layer performs skip connection fusion on the current upsampled feature map and the corresponding scale feature map from the encoder. To improve the fusion quality, the encoder feature map has been weighted and enhanced through a channel attention module. The fusion method uses channel dimension concatenation (Concat), followed by a 1×1 convolution for channel compression and non-linear mapping to ensure consistent feature dimensions and provide a unified input channel number for the subsequent residual convolution module.

[0044] The double-layer residual convolution layer extracts local spatial context information, refines the object boundary, and enhances the feature robustness of the fused feature map.

[0045] In the present invention, the 4 decoder units adopt exactly the same structure, as Figure 3 shown. Each decoder unit consists of a fusion layer, a convolution layer, and a double-layer residual convolution layer; one input of the fusion layer forms the main path input of the decoder unit, and the other input of the fusion layer forms the skip input of the decoder unit. The output of the fusion layer is connected to the input of the convolution layer, the output of the convolution layer is connected to the input of the double-layer residual convolution layer, and the output of the double-layer residual convolution layer forms the output of the decoder unit.

[0046] (5) Downsampling and Upsampling

[0047] Max pooling (2×2, stride 2) is used for downsampling to reduce the number of pixels while retaining key information. Transposed convolution (kernel size 3×3, stride 2) is used for upsampling to gradually restore the resolution.

[0048] (6) Output Layer

[0049] The output layer consists of a bilinear interpolation amplification layer, a 1×1 convolution layer, and a Sigmoid activation function layer; the input of the bilinear interpolation amplification layer forms the input of the output layer, the output of the bilinear interpolation amplification layer is connected to the input of the 1×1 convolution layer, the output of the 1×1 convolution layer is connected to the input of the Sigmoid activation function layer, and the output of the Sigmoid activation function layer forms the output of the output layer. The output layer generates a pixel-level tumor segmentation mask map.

[0050] Step 2: Obtain the breast DCE-MRI dataset, and use the extended Tofts model to extract the breast PK feature maps corresponding to each group of breast DCE-MRI temporal image sequences in the breast DCE-MRI dataset, and obtain the PK feature map dataset of the breast.

[0051] In the present invention, the breast DCE-MRI dataset is derived from the publicly available dataset BreastDM and the clinical breast DCE-MRI of the collaborating hospitals. Each group of breast DCE-MRI temporal images in this dataset is the temporal image of breast tumor DCE-MRI with tumors. The image format of DCE-MRI is DICOM, which includes multi-phase dynamic contrast-enhanced sequences. The image data of each patient includes a baseline (non-enhanced phase) and multiple enhanced phases, and the phase characteristics are as Figure 4 shown.

[0052] First, preprocess each group of DCE-MRI temporal images in the breast DCE-MRI dataset, including registration, normalization, and noise suppression. Rigid registration algorithm based on mutual information is used for registration to eliminate displacement errors caused by patient breathing or movement. Normalization normalizes the pixel values of each frame of the image by Z-Score9. Non-local means filter is applied for noise suppression to remove high-frequency noise and retain the details of the tumor edge.

[0053] Then, use the extended Tofts model to extract the breast PK feature maps corresponding to each group of breast DCE-MRI temporal images in the breast DCE-MRI dataset. In DCE-MRI imaging medicine, PK parameters are a set of quantitative indicators obtained through contrast agent kinetic analysis, including the transfer rate constant K trans , blood volume fraction v p , extracellular space fraction v e , which can reflect the angiogenesis characteristics and microcirculation perfusion of tumors. PK parameters can be extracted using the extended Tofts model (Extended Tofts Model, ETM). The extended Tofts model is a commonly used semi-quantitative model in pharmacokinetic modeling, which is widely used in DCE-MRI to analyze the uptake and clearance processes of tissues to contrast agents, so as to evaluate microenvironment characteristics such as tumor blood flow, vascular permeability, and extravascular space volume. The extended Tofts model can be described by the following formula:

[0054]

[0055] where, C t (t) is the total contrast agent concentration in the tissue; C p (*) is the contrast agent concentration in the plasma (arterial input function, AIF); V p is the intravascular plasma volume fraction; K trans is the vascular permeability parameter, indicating the rate of contrast agent entering the extravascular space from the intravascular; K epIt is used to describe the rate at which the contrast agent returns from the extracellular extravascular space to the plasma. V e is the volume fraction of the extravascular space.

[0056] Specifically, the process of extracting the PK feature map using the extended Tofts model is as follows:

[0057] First, the time-signal intensity curve S(t) of each voxel is extracted from the DCE-MRI temporal image, as Figure 5 shown, and the signal intensity is converted to the contrast agent concentration according to the following formula, that is, the total contrast agent concentration C t (t) in the tissue:

[0058]

[0059] where r1 is the transverse relaxation rate of the contrast agent (usually taken as 4.5 s-1·mM-1); T 10 is the T1 value of the tissue before injection (using the reference average value of 1000 ms); S(t) is the DCE-MRI signal intensity at a certain time sequence; S0 is the signal intensity before injection (pre-enhancement image).

[0060] Then, the Parker AIF model (population-averaged AIF model) is used to obtain the arterial input function, that is, the contrast agent concentration C p (t) in the plasma.

[0061] Next, the total contrast agent concentration C t (t) in the tissue and the contrast agent concentration C p (t) in the plasma are fed into the extended Tofts model, and the PK parameters, that is, the transport rate constant K trans , blood volume fraction v p , extracellular space fraction v e are solved using the TRR algorithm (Trust Region Reflective, an iterative optimization algorithm for nonlinear least squares problems).

[0062] Finally, the PK parameter feature map is generated by fitting using the PK parameters, that is, the transport rate constant K trans , blood volume fraction v p , extracellular space fraction v e , as Figure 6 shown.

[0063] Step 3: Use the breast DCE-MRI dataset and the breast PK feature map dataset to train and validate the spatio-temporal fusion network to obtain an accurate breast tumor segmentation model. Among them, the breast DCE-MRI dataset is input from the main path of the spatio-temporal fusion network, and the breast PK feature map dataset is input from the auxiliary input of the spatio-temporal fusion network.

[0064] During the training process, the two datasets of breast DCE-MRI dataset and breast PK feature map dataset are respectively divided into a training set (70%), a validation set (15%) and a test set (15%) according to a certain proportion. The training process is implemented using the PyTorch framework. The optimizer is selected as Adam, the initial learning rate is set to 0.001, and the cosine annealing learning rate scheduling strategy is adopted. The training is iterated for a total of 100 epochs. The input image size for each batch is B×T×1×H×W, where the batch size B = 16, and T is the number of time phases; the size of the PK feature map corresponding to each sample is B×T×3×H×W.

[0065] During the validation process, quantitative evaluation is carried out first. The quantitative indicators include the Dice coefficient (DSC) and the intersection over union (IoU). Among them, the Dice coefficient evaluates the segmentation accuracy, and the intersection over union measures the coverage of the tumor area by the model. Then qualitative evaluation is carried out. Using the tumor area marked by experts as a comparison, the segmentation results are visually inspected.

[0066] Step 4: Obtain the breast DCE-MRI to be segmented, and use the extended Tofts model to extract the corresponding breast PK feature map of the breast DCE-MRI to be segmented.

[0067] Step 5: Send the breast DCE-MRI to be segmented and the breast PK feature map to be segmented into the breast tumor precise segmentation model to obtain the segmentation results. Among them, the breast DCE-MRI to be segmented is input from the main path of the breast tumor precise segmentation model, and the breast PK feature map to be segmented is input from the auxiliary input of the breast tumor precise segmentation model.

[0068] The breast DCE-MRI to be segmented can be the breast DCE-MRI with tumors or the breast DCE-MRI without tumors. After the breast DCE-MRI to be segmented and the breast PK feature map to be segmented are sent into the breast tumor precise segmentation model to obtain the segmentation results, post-processing can be further carried out on the segmentation results, including CRF boundary optimization, artifact removal and result quantification. For CRF boundary optimization, the DenseCRF library is used, and the parameter settings are: the Gaussian kernel scale is 5 pixels; the window range is 9×9. Artifact removal removes isolated small regions, and noise artifacts are filtered by the area threshold. Result quantification calculates indicators such as tumor volume, enhancement intensity, and boundary irregularity according to the segmentation results, and generates a diagnostic report.

[0069] The present invention combines deep learning and pharmacokinetic modeling, and utilizes the deep learning of the spatio-temporal dynamic information and radiomics features of DCE-MRI, aiming to improve the segmentation accuracy and medical interpretability of breast tumors, and enhance clinical usability, enabling doctors to make more accurate diagnoses and treatment plans based on the segmentation results.

[0070] It should be noted that although the embodiments described above of the present invention are illustrative, they are not limitations on the present invention. Therefore, the present invention is not limited to the above specific embodiments. Without departing from the principle of the present invention, any other embodiments obtained by those skilled in the art under the inspiration of the present invention are deemed to be within the protection scope of the present invention.

Claims

1. An accurate segmentation method for breast tumors based on a spatio-temporal fusion network, characterized in that, It includes the following steps: Step 1: Construct a spatio-temporal fusion network based on the Unet structure; The spatio-temporal fusion network consists of an input layer, 4 encoder units, 3 downsampling units, 3 attention skip connection units, 4 upsampling units, 4 decoder units, and an output layer; the input of the input layer forms the main path input of the spatio-temporal fusion network, and the auxiliary inputs of the first encoder unit, the second encoder unit, the third encoder unit, and the fourth encoder unit form the auxiliary input of the spatio-temporal fusion network; the output of the input layer is connected to the main path input of the first encoder unit, the output of the first encoder unit is connected to the main path input of the second encoder unit through the first downsampling unit, the output of the second encoder unit is connected to the main path input of the third encoder unit through the second downsampling unit, and the output of the third encoder unit is connected to the main path input of the fourth encoder unit through the third downsampling unit; the output of the fourth encoder unit is connected to the main path input of the first decoder unit through the first upsampling unit, the output of the first decoder unit is connected to the main path input of the second decoder unit through the second upsampling unit, the output of the second decoder unit is connected to the main path input of the third decoder unit through the third upsampling unit, and the output of the third decoder unit is connected to the main path input of the fourth decoder unit through the fourth upsampling unit; the output of the first encoder unit is also connected to the skip input of the third decoder unit through the first attention skip connection unit, the output of the second encoder unit is also connected to the skip input of the second decoder unit through the second attention skip connection unit, the output of the third encoder unit is also connected to the skip input of the first decoder unit through the third attention skip connection unit, the output of the fourth decoder unit is connected to the input of the output layer, and the output of the output layer forms the output of the spatio-temporal fusion network; Step 2: Obtain the breast DCE-MRI dataset, and use the extended Tofts model to extract the breast PK feature maps corresponding to each group of breast DCE-MRI temporal image maps in the breast DCE-MRI dataset to obtain the breast PK feature map dataset; Step 3: Use the breast DCE-MRI dataset and the breast PK feature map dataset to train and validate the spatio-temporal fusion network to obtain an accurate breast tumor segmentation model; Step 4: Obtain the breast DCE-MRI to be segmented, and use the extended Tofts model to extract the breast PK feature map to be segmented corresponding to the breast DCE-MRI to be segmented; Step 5: Send the breast DCE-MRI to be segmented and the breast PK feature map to be segmented into the accurate breast tumor segmentation model to obtain the segmentation result.

2. The accurate breast tumor segmentation method based on a spatio-temporal fusion network according to claim 1, characterized in that The first encoder unit and the fourth encoder unit are both composed of 3 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 3 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the third residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit. The second encoder unit is composed of 4 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 4 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the fourth residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit. The third encoder unit is composed of 6 residual convolutional layers, an LSTM, a fusion layer, and a 3D convolutional layer; the 6 residual convolutional layers are connected in series in sequence. The input of the first residual convolutional layer forms the main path input of this encoder unit. The output of the sixth residual convolutional layer is connected to the input of the LSTM. The output of the LSTM is connected to one input of the fusion layer. The other input of the fusion layer forms the auxiliary input of this encoder unit. The output of the fusion layer forms the output of this encoder unit.

3. A precise segmentation method for breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that, Each decoder unit is composed of a fusion layer, a convolutional layer, and a double-layer residual convolutional layer; one input of the fusion layer forms the main path input of this decoder unit. The other input of the fusion layer forms the skip input of this decoder unit. The output of the fusion layer is connected to the input of the convolutional layer. The output of the convolutional layer is connected to the input of the double-layer residual convolutional layer. The output of the double-layer residual convolutional layer forms the output of this decoder unit.

4. A precise segmentation method for breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that The attention skip connection unit is a SE channel attention unit.

5. A precise segmentation method for breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that, The input layer is composed of a 7×7 convolutional layer and a 3×3 pooling layer; the input of the 7×7 convolutional layer forms the input of the input layer. The output of the 7×7 convolutional layer is connected to the input of the 3×3 pooling layer. The output of the 3×3 pooling layer forms the output of the input layer.

6. The precise segmentation method of breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that, The output layer is composed of a bilinear interpolation upsampling layer, a 1×1 convolutional layer, and a Sigmoid activation function layer; the input of the bilinear interpolation upsampling layer forms the input of the output layer. The output of the bilinear interpolation upsampling layer is connected to the input of the 1×1 convolutional layer. The output of the 1×1 convolutional layer is connected to the input of the Sigmoid activation function layer. The output of the Sigmoid activation function layer forms the output of the output layer.

7. A precise segmentation method for breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that In step 3, the breast DCE-MRI dataset is fed into the main path input of the spatio-temporal fusion network, and the breast PK feature map dataset is fed into the auxiliary input of the spatio-temporal fusion network.

8. A precise segmentation method for breast tumors based on a spatio-temporal fusion network according to claim 1, characterized in that, In step 5, the breast DCE-MRI to be segmented is fed into the main path input of the breast tumor precise segmentation model, and the breast PK feature map to be segmented is fed into the auxiliary input of the breast tumor precise segmentation model.

Citation Information

Cited By

  • Pharmacokinetic parameter estimation method and device, system and storage medium

    CN122049598A

  • Pharmacokinetic parameter estimation method and device, system, storage medium

    CN122049598B