BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning
Patent Information
- Application Number
- CN202311357406.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-10-18
AI Technical Summary
然而,人工读取MRI非常耗时,计算机辅助早期诊断乳腺病变具有重要意义
[0055]本发明提出了一种基于协同多任务学习的BI-RADS四类乳腺病变分割与分类方法。模型包含三个子网,分别是初步分割子网、分类子网和重分割子网。给定一个乳腺的MRI图像Ii,Ii∈H×W×C,图像级标签ls∈l1,…,lN和像素级标签Yseg∈H×W,其中H和W分别为MRI图像的高度和宽度,C是通道数,N是类别数;首先,MRI图像Ii通过初步分割子网中的三个流解码器勾画出病变的分割图片和边界分割图
在区分良恶性病变时,分类子网将MRI图像Ii、
与Ii通道连接后的图像
与Ii通道连接后的图像
作为输入,分类的工作方式是通过分割结果来促进分类结果;然后,利用重分割子网对分类结果进行细化,得到最终的分割结果;通过三个子网的协作,分割和分类任务相互促进,提高性能。
Smart Images

Figure CN117372766B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of breast magnetic resonance imaging segmentation and classification technology, specifically to a BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning. Background Technology
[0002] Currently, the Breast Imaging Reporting and Data System (BI-RADS) grading relies on ultrasound technology. The four categories of breast lesions are divided into 4A, 4B, and 4C, reflecting an increased likelihood of low-grade malignancy (2%-10%), intermediate-grade malignancy (10%-50%), and high-grade malignancy (50%-95%), respectively. Because there is some overlap between benign and malignant lesions in the four BI-RADS categories, the probability of a lesion being malignant ranges from 2% to 95%. This leads to some benign lesions being over-evaluated and subjected to invasive treatment, while some malignant lesions are underestimated, delaying optimal treatment. Therefore, clearly defining the benign or malignant nature of breast lesions in the four BI-RADS categories is crucial for accurate diagnosis and follow-up treatment.
[0003] Breast magnetic resonance imaging (MRI) has broken away from the traditional single-morphology diagnostic model, allowing for quantitative and qualitative analysis of lesions. However, manually interpreting MRI images is very time-consuming, making computer-aided early diagnosis of breast lesions of great significance.
[0004] However, MRI computer diagnosis has the following problems: (1) lesions are usually irregular in shape and not fixed in location; (2) lesions only account for 0.01%-4.63% of the entire image, resulting in serious class imbalance; (3) lesions have low contrast with normal tissue, making them difficult to identify. Therefore, completing this task is very challenging.
[0005] Deep learning is one of the most effective methods for medical image analysis. However, due to inherent problems with breast MRI images, there is still significant room for improvement in their diagnostic accuracy. Much work has been dedicated to addressing these issues. To address problems arising from irregular shapes and non-fixed locations, some existing techniques attempt to capture multi-scale information from images using kernels of different sizes and resolutions. To address class imbalance, some existing techniques enhance the recognition of small targets through attention mechanisms. To address the problem of low contrast between lesions and normal tissue, some existing techniques utilize boundary-aware mechanisms to characterize blurred lesion boundaries. These methods have been successfully applied to organ and lesion image analysis. However, they cannot be directly applied to the reading of breast MRI images.
[0006] This is primarily because breast lesions only account for 0.01%–4.63% of the image. Existing methods for addressing class imbalance are only effective when lesions occupy more than 10% of the image. When reading breast MRI images, these methods often treat lesions as noise. This exacerbates the problems of irregular lesion shapes, variable locations, and low contrast.
[0007] Clinically, doctors first determine the location of the lesion and then perform a pathological examination to differentiate between benign and malignant lesions. Furthermore, benign or malignant lesions can manifest in different forms. For example, in the case of breast lesions, round or oval lesions are usually benign, while irregularly shaped lesions may be malignant. This illustrates that the segmentation result is closely related to the accuracy of benign / malignant differentiation. Inspired by this, a method that simultaneously considers segmentation and classification tasks to mutually reinforce each other is a promising solution for achieving positive results. Summary of the Invention
[0008] To address the aforementioned technical problems, this invention provides a BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning.
[0009] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0010] A BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning is proposed. This method uses a collaborative multi-task learning segmentation and classification network to segment BI-RADS category 4 breast lesions in MRI images, obtaining the segmentation results. Specifically, the method includes the following steps:
[0011] Step 1: Preprocess the acquired MRI images and divide the preprocessed MRI images into training dataset and test dataset;
[0012] Step 2: Construct the segmentation and classification network; the segmentation and classification network includes an encoder, a preliminary segmentation subnet, a classification subnet, and a re-segmentation subnet;
[0013] The MRI image is input into an encoder based on the ResNet-34 structure to obtain five layers of image features, which are denoted as the first layer image feature χ1, the second layer image feature χ2, the third layer image feature χ3, the fourth layer image feature χ4, and the fifth layer image feature X5, respectively.
[0014] The initial segmentation subnet consists of three parallel stream decoders: a location stream decoder, a segmentation stream decoder, and a boundary stream decoder. The inputs to the location stream decoder are X4 and X5, the inputs to the segmentation stream decoder are X2, X3, X4, and X5, and the inputs to the boundary stream decoder are X1, X2, X3, and X4. The location stream decoder is used to extract features within the lesion to obtain the location segmentation map. The boundary stream decoder is used to refine the boundaries of objects and output a boundary segmentation map. The segmentation stream decoder generates segmented images by acquiring position and boundary information. By integrating multi-scale receptive fields to adapt to lesions of different sizes;
[0015] MRI image I i , with I i Image after channel concatenation with I i Image after channel concatenation The modal features are fed into the three branches of the classification subnet, each branch of which consists of a backbone network composed of a ResNet-18 model and a fully connected layer after the backbone network. The modal features output by the fully connected layers of the three branches are connected to obtain multimodal features. The multimodal features are then input into a pre-trained classifier to obtain the classification results. The classification results include benign and malignant results.
[0016] The resegmentation subnet uses the benign and malignant results given by the segmentation subnet to refine the segmentation results: based on the different characteristics of benign and malignant results, two convolutional layers with activation functions are created. Benign sample I2 is fed into the benign classification head, and malignant sample I1 is fed into the malignant classification head to obtain the final segmentation result.
[0017] Step 3: Set loss functions for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network, and sum the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder to obtain the loss function of the initial segmentation network; train the initial segmentation network using the training dataset;
[0018] Loss functions are set for the classification subnet and the resegmentation subnet, and the loss functions of the classification subnet and the resegmentation subnet are added together to obtain the loss function of the entire segmentation and classification network. The segmentation and classification networks are trained using the loss functions of the segmentation and classification networks and the training dataset.
[0019] Step four: Input the MRI images from the test dataset into the trained segmentation and classification network to obtain the segmentation results.
[0020] Furthermore, the location stream decoder is used to extract features within the lesion to obtain a location segmentation map. Specifically, this includes: using χ4 and χ5 in the location stream decoder to extract internal features of the lesion; transposing and convolving χ5 and mapping it to a space of the same size as χ4; and then concatenating χ5 and χ4 to obtain χ. spatial Then χ spatialThe input is fed into the attention block to obtain the correct location of the lesion. The formula for the attention block is as follows:
[0021]
[0022] in, ψ(·) represents element-wise multiplication; σ(·) represents convolution; χ represents non-linear activation; spatial It is the result of splicing χ4 and χ5, α inner It is the channel scale factor.
[0023] Furthermore, the boundary flow decoder is used to refine the boundaries of objects, outputting a boundary segmentation map. Specifically, this includes:
[0024] Lesion boundaries are predicted using feature maps with clear boundaries χ1, χ2, χ3, and χ4. Three gated convolutional layers are added to disable irrelevant features extracted by the Res encoder; S l Defined as the feature map output by the l-th layer boundary flow decoder, and the attention map α of the l-th layer gated convolutional layer. l for:
[0025] α1=σ(Ψ(S1||Ψ(χ1)));
[0026] Where ψ(·) represents convolution operation; σ(·) represents nonlinear activation; || represents channel concatenation of feature maps; δ l The residual block is composed of two normalized convolutional layers with skip connections, resulting in the feature map S of the l+1 layer boundary flow decoder. l+1 :
[0027]
[0028] Where Res(·) represents the residual block, and the original χ3 is replaced by the result of connecting Res(χ1) and χ3; the output of the last layer boundary stream decoder is the boundary segmentation map.
[0029] Furthermore, the segmentation stream decoder generates segmented images by acquiring location and boundary information. Specifically, the process includes: the segmentation stream decoder uses the UNet structure to decode χ2, χ3, χ4, and χ5, and obtains feature maps through convolution, pooling, and activation operations of dense dilated convolutions, which are then input to the next layer to finally obtain the segmented image.
[0030] Furthermore, in step one, the preprocessing of the acquired MRI images specifically includes:
[0031] Preprocessing of MRI images: Rotating MRI images at angles from 30° to 90°, and randomly flipping MRI images horizontally and vertically; z-score normalization of MRI images using mean μ = 0 and standard deviation σ = 1.
[0032] Further, in step three, loss functions are set for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network. When the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder are added together to obtain the loss function of the initial segmentation network, the specific steps include:
[0033] Loss function L of the position stream decoder inner for:
[0034]
[0035] Among them, Y inner for Corresponding tags;
[0036] The loss function L of the boundary stream decoder edge for:
[0037]
[0038] Among them, Y edge for Corresponding tags;
[0039] Loss function L of the split stream decoder seg for:
[0040]
[0041] Among them, Y seg for Corresponding tags;
[0042] The loss function L of the initial segmentation network init for:
[0043] L init =L seg +L edge +L inner .
[0044] Furthermore, when setting loss functions for the classification subnet and the resegmentation subnet, and adding the loss functions of the classification subnet and the resegmentation subnet to obtain the loss function of the entire segmentation and classification network, the specific steps include:
[0045] The classification subnetwork uses the original MRI image, along with the obtained location segmentation map, boundary segmentation map, and classification results, and applies binary cross-entropy as the loss to obtain the loss function L of the classification subnetwork. cls :
[0046] L cls =∑λ mod L mod ;
[0047] Where λ mod The weights of the loss function for each branch of the classification subnet are represented by ∑λ. mod =1,L mod This represents the loss function for each branch of the classification subnet;
[0048] The reclassified subnet will use binary cross-entropy loss to optimize the final result, resulting in the loss function L in the reclassified subnet. final :
[0049] L final =(Y seg -1)log(1-M)-Y seg logM;
[0050] M represents the final segmentation result obtained by re-segmenting the subnet; Y seg for Corresponding tags;
[0051] The loss function L of the entire segmentation and classification network is:
[0052] L=λ1L final +λ2L cls ;
[0053] λ1 and λ2 are weight parameters.
[0054] Compared with the prior art, the beneficial technical effects of the present invention are:
[0055] This invention proposes a BI-RADS four-class breast lesion segmentation and classification method based on collaborative multi-task learning. The model consists of three subnetworks: a preliminary segmentation subnetwork, a classification subnetwork, and a resegmentation subnetwork. Given a breast MRI image I i I i ∈H×W×C, image-level label l s ∈l1,…,l N and pixel-level label Y seg ∈H×W, where H and W are the height and width of the MRI image, respectively, C is the number of channels, and N is the number of categories; first, the MRI image I i The lesion segmentation images were delineated using three stream decoders in the initial segmentation subnet. and boundary segmentation map When differentiating between benign and malignant lesions, the classification subnetwork will use MRI images I i , with I i Image after channel concatenation with I i Image after channel concatenation As input, the classification process works by using segmentation results to improve classification results; then, the classification results are refined using a resegmentation subnet to obtain the final segmentation result; through the collaboration of the three subnets, the segmentation and classification tasks mutually promote each other, improving performance. Attached Figure Description
[0056] Figure 1 This is a schematic diagram of the segmentation and classification network of the present invention;
[0057] Figure 2 This is a schematic diagram of the initial subnetting of the present invention. Detailed Implementation
[0058] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0059] This invention designs a novel network for reading BI-RADS Category 4 MRI breast lesions, namely a collaborative multi-task learning segmentation and classification network. First, a preliminary segmentation subnetwork with a decoder of three input streams is constructed for initial segmentation, employing dense dilated convolution, attention mechanisms, and upsampling boundary gates. Second, a classification subnetwork is constructed to distinguish between benign and malignant lesions. This classification subnetwork comprehensively considers lesion masks, boundary masks, and original image information, improving classification accuracy. Finally, a re-segmentation subnetwork is constructed to further enhance segmentation accuracy. The re-segmentation subnetwork uses the classification results generated by the classification subnetwork and creates two pixel-level conditional classifiers to predict the masks for benign and malignant lesions, respectively. The preliminary segmentation subnetwork, classification subnetwork, and re-segmentation subnetwork work collaboratively to provide lesion segmentation and classification results.
[0060] Specifically, the following steps are included:
[0061] Step 1: Preprocess the acquired MRI images and divide the preprocessed MRI images into training dataset and test dataset;
[0062] This invention collected MRI image data from patients at a hospital. The inclusion criteria for this dataset were threefold: (i) breast MRI examination performed within one week prior to surgery, diagnosed as BI-RADS category IV lesions; (ii) postoperative pathological examination of the lesions, obtaining pathological results; and (iii) exclusion of patients who had undergone surgery (radiotherapy, chemotherapy, biopsy, etc.) prior to the MRI examination, affecting the morphology or nature of the lesions. A total of 248 patients met the inclusion criteria, including 75 benign cases and 173 malignant cases. The MRI images of all 248 patients had lesion segmentation markers and benign / malignant classification markers. This dataset will help train a model that effectively assists physicians in diagnosing BI-RADS category IV lesions.
[0063] The dataset was preprocessed before being input into the CMTL-SC model. The 248 samples were divided into training and testing subsets, with 70% used for training and 30% for testing. The image size ranged from 384×384 to 768×768. For each image, the size was resized to 512×512. To prevent data imbalance from degrading model performance, the images were augmented twice. Augmentation included rotations from 30 to 90 degrees and random horizontal and vertical flips. Finally, z-score normalization was performed using a mean μ = 0 and a standard deviation σ = 1. It is important to note that for the normalization of subtractive images, the subtraction was obtained first, followed by intensity normalization.
[0064] Step 2: Construct the segmentation and classification network; the segmentation and classification network includes an encoder, a preliminary segmentation subnet, a classification subnet, and a re-segmentation subnet.
[0065] The MRI image is input to an encoder based on the ResNet-34 structure to obtain five layers of image features, which are denoted as the first layer image feature X1, the second layer image feature χ2, the third layer image feature χ3, the fourth layer image feature X4, and the fifth layer image feature X5.
[0066] The initial segmentation subnet consists of three parallel stream decoders: a location stream decoder, a segmentation stream decoder, and a boundary stream decoder. The inputs to the location stream decoder are X4 and X5, the inputs to the segmentation stream decoder are X2, X3, X4, and X5, and the inputs to the boundary stream decoder are X1, X2, X3, and X4. The location stream decoder is used to extract features within the lesion to obtain the location segmentation map. The boundary stream decoder is used to refine the boundaries of objects and output a boundary segmentation map. The segmentation stream decoder generates segmented images by acquiring position and boundary information. By integrating multi-scale receptive fields, it adapts to lesions of different sizes.
[0067] Specifically, the location stream decoder is used to extract features inside the lesion to obtain a location segmentation map. Specifically, this includes: using χ4 and χ5 in the location stream decoder to extract internal features of the lesion; transposing and convolving χ5 and mapping it to a space of the same size as χ4; and then concatenating χ5 and χ4 to obtain χ. spatial Then χ spatial The input is fed into the attention block to obtain the correct location of the lesion. The formula for the attention block is as follows:
[0068]
[0069] in, ψ(·) represents element-wise multiplication; σ(·) represents convolution; χ represents non-linear activation; spatial It is the result of splicing χ4 and χ5, α inner It is the channel scale factor.
[0070] Specifically, the boundary stream decoder is used to refine the boundaries of objects and output a boundary segmentation map. Specifically, this includes:
[0071] Lesion boundaries are predicted using feature maps with clear boundaries χ1, χ2, χ3, and χ4. Three gated convolutional layers are added to disable irrelevant features extracted by the Res encoder; S l Defined as the feature map output by the l-th layer boundary flow decoder, and the attention map α of the l-th layer gated convolutional layer. l for:
[0072] α1=σ(Ψ(S1||Ψ(χ1)));
[0073] Where ψ(·) represents convolution operation; σ(·) represents nonlinear activation; || represents channel concatenation of feature maps; δ l The residual block is composed of two normalized convolutional layers with skip connections, resulting in the feature map S of the l+1 layer boundary flow decoder. l+1 :
[0074]
[0075] Where Res(·) represents the residual block, and the original χ3 is replaced by the result of connecting Res(χ1) and χ3; the output of the last layer boundary stream decoder is the boundary segmentation map.
[0076] Specifically, the segmentation stream decoder generates segmented images by acquiring positional and boundary information. Specifically, the process includes: the segmentation stream decoder uses the UNet structure to decode χ2, χ3, χ4, and χ5, and obtains feature maps through convolution, pooling, and activation operations of dense dilated convolutions, which are then input to the next layer to finally obtain the segmented image.
[0077] In the preceding text, to avoid gradient vanishing, the original image was fed into a ResNet-34 structure, which served as the encoder. The encoder retained the first four feature extraction blocks of the ResNet-34 structure. The average pooling layer before the fully connected layers was removed. Due to the imbalance in the distribution between lesions and the background, segmentation using only the UNet structure would result in the disappearance of small lesions. Therefore, a positional flow was designed to allow the network to focus on feature extraction within the lesion, preventing the loss of lesion information. Downsampling weakens or eliminates image edges. Intuitively, if downsampling results in information loss, upsampling can mitigate this and help discover hidden structures. Therefore, a boundary flow was designed to refine the boundary information of objects by adding an upsampling layer. Finally, a segmentation flow was designed to generate segmentation results by acquiring positional and boundary information, and to adapt to lesions of different sizes by integrating multi-scale receptive fields. The role of the positional flow: Higher-level features have more semantic information, which helps in object localization. The role of the boundary flow: Lower-level features originate from shallow networks and have rich spatial information, enabling the network to focus more on lesion boundary extraction.
[0078] MRI image I i , with I i Image after channel concatenation with I i Image after channel concatenation The modal features are fed into three branches of the classification subnetwork. Each branch consists of a backbone network based on a ResNet-18 model and a fully connected layer following the backbone. The modal features output from the fully connected layers of the three branches are concatenated to obtain multimodal features. These multimodal features are then input into a pre-trained classifier to obtain classification results, including benign and malignant results. The classification subnetwork replicates the ResNet-18 base model three times, constructing a weight-sharing mechanism to extract different features and suppress overfitting. At the end of each branch, a fully connected layer is used to classify each modality. After generating specific modal features, they are concatenated to obtain multimodal features, and the classifier is trained for final prediction. The multimodal features are designed to encourage competition among the three modalities.
[0079] The re-segmentation subnet refines the segmentation results using the benign and malignant results provided by the initial segmentation subnet. Based on the different features of benign and malignant results, two convolutional layers with activation functions are created. Benign sample I2 is fed into the benign classification head, and malignant sample I1 is fed into the malignant classification head, yielding the final segmentation result. In the initial segmentation subnet, the class head in the segmentation stream aggregates all features. The re-segmentation subnet attempts to find a global class center to identify all variations between samples. This invention refers to this global class center as the global class head. However, the global class head often incorrectly identifies pixels that look different but belong to the same category as different categories. Since pixels of the same class have sample-specific class centers, they are easier to identify than global class centers. Therefore, this invention uses the benign and malignant results provided by the initial segmentation subnet to refine the segmentation results. In the re-segmentation subnet, the class head of the segmentation stream is retrained. Based on the different features of benign and malignant lesions, two 3×3 convolutional layers with activation functions are created. It allows two classifiers to focus on specific sample differences for each class, sending benign samples into the benign class head and malignant samples into the malignant class head.
[0080] Step 3: Set loss functions for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network, and sum the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder to obtain the loss function of the initial segmentation network; train the initial segmentation network using the training dataset;
[0081] Loss functions are set for the classification subnet and the resegmentation subnet, and the loss functions of the classification subnet and the resegmentation subnet are added together to obtain the loss function of the entire segmentation and classification network. The segmentation and classification networks are trained using the loss functions of the segmentation and classification networks and the training dataset.
[0082] Specifically, in step three, loss functions are set for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network. When the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder are added together to obtain the loss function of the initial segmentation network, the following steps are taken:
[0083] Loss function L of the position stream decoder inner for:
[0084]
[0085] Among them, Y inner for Corresponding tags;
[0086] The loss function L of the boundary stream decoder edge for:
[0087]
[0088] Among them, Y edge for Corresponding tags;
[0089] Loss function L of split stream decoder seg for:
[0090]
[0091] Among them, Y seg for Corresponding tags;
[0092] The loss function L of the initial segmentation network init for:
[0093] L init =L seg +L edge +L inner .
[0094] Specifically, when setting loss functions for the classification subnet and the resegmentation subnet, and adding the loss functions of the classification subnet and the resegmentation subnet to obtain the loss function of the entire segmentation and classification network, the specific steps include:
[0095] The classification subnetwork uses the original MRI image, along with the obtained location segmentation map, boundary segmentation map, and classification results, and applies binary cross-entropy as the loss to obtain the loss function L of the classification subnetwork. cls :
[0096] L cls =∑λ mod L mod ;
[0097] Where λ mod The weights of the loss function for each branch of the classification subnet are represented by ∑λ. mod =1,L mod This represents the loss function for each branch of the classification subnet;
[0098] The reclassified subnet will use binary cross-entropy loss to optimize the final result, resulting in the loss function L in the reclassified subnet. final :
[0099] L final =(Y seg -1)log(1-M)-Y seg logM;
[0100] M represents the final segmentation result obtained by re-segmenting the subnet; Y seg for Corresponding tags;
[0101] The loss function L of the entire segmentation and classification network is:
[0102] L=λ1L final +λ2L cls ;
[0103] λ1 and λ2 are weight parameters.
[0104] Step four: Input the MRI images from the test dataset into the trained segmentation and classification network to obtain the segmentation results.
[0105] To evaluate the segmentation performance of the segmentation and classification networks, this invention employs five metrics: dice similarity coefficient, Jaccard index, pixel sensitivity, pixel specificity, and 95% Hausdorff distance. To evaluate classification performance, this invention employs five additional metrics: accuracy, sensitivity, specificity, area under the acceptance operator curve, and F1 score. This invention qualitatively and quantitatively evaluates the lesion segmentation performance of the model. The results of the proposed method and other segmentation methods are shown in Table 1. Experimental results show that the segmentation network proposed in this invention achieves significant results in the segmentation of four types of breast lesions in BI-RADS. Specifically, the method of this invention yields an average dice similarity coefficient of 70.4%, a Jaccard index of 54.9%, pixel sensitivity and pixel specificity of 67.5% and 76.8%, respectively, and a 5% Hausdorff distance of 13.8. Among all the segmentation comparison methods, the method of this invention produces the best results for most metrics. Experimental results are shown in Table 1.
[0106] Table 1
[0107]
[0108]
[0109] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0110] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning, which segments BI-RADS category 4 breast lesions in MRI images through a segmentation and classification network based on collaborative multi-task learning, and obtains the segmentation results; Specifically, the following steps are included: Step 1: Preprocess the acquired MRI images and divide the preprocessed MRI images into training dataset and test dataset; Step 2: Construct the segmentation and classification network; the segmentation and classification network includes an encoder, a preliminary segmentation subnet, a classification subnet, and a re-segmentation subnet; MRI images are input into a ResNet-34-based encoder to obtain five layers of image features, denoted as the first layer of image features. Second layer image features Third layer image features Fourth layer image features Fifth layer image features ; The initial subnetting consists of three parallel stream decoders: a positional stream decoder, a segmented stream decoder, and a boundary stream decoder; the input to the positional stream decoder is... , The input to the split stream decoder is , , , The input to the boundary stream decoder is , , , The location stream decoder is used to extract features inside the lesion to obtain a location segmentation map. The boundary flow decoder is used to refine the boundaries of objects and output a boundary segmentation map. The segmentation stream decoder generates segmented images by acquiring location and boundary information. It adapts to lesions of different sizes by integrating multi-scale receptive fields; MRI images , and Image after channel concatenation , and Image after channel concatenation The modal features are input into the three branches of the classification subnetwork, each branch of which includes a backbone network composed of a ResNet-18 model and a fully connected layer after the backbone network. The modal features output from the fully connected layers of the three branches are connected to obtain multimodal features. The multimodal features are then input into a pre-trained classifier to obtain the classification results. The classification results include benign and malignant results. The resegmentation subnetwork refines the segmentation results using the benign and malignant results provided by the previous segmentation subnetwork: based on the different characteristics of benign and malignant results, two convolutional layers with activation functions are created to process the benign samples... Send the samples into the benign classification head to separate the malignant samples. The data is fed into the malignant classification head to obtain the final segmentation result; Step 3: Set loss functions for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network, and sum the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder to obtain the loss function of the initial segmentation network; train the initial segmentation network using the training dataset; Loss functions are set for the classification subnet and the resegmentation subnet, and the loss functions of the classification subnet and the resegmentation subnet are added together to obtain the loss function of the entire segmentation and classification network; the segmentation and classification networks are trained using the loss functions of the segmentation and classification networks and the training dataset; Step 4: Input the MRI images from the test dataset into the trained segmentation and classification network to obtain the segmentation results; The location stream decoder is used to extract features inside the lesion to obtain a location segmentation map. Specifically, this includes using it in the location stream decoder. , To extract internal features of the lesion, Perform a transposed convolution and map it to... In a space of the same size, then and By splicing them together, we get Then The input is fed into the attention block to obtain the correct location of the lesion. The formula for the attention block is as follows: ; in, This indicates element-wise multiplication; This represents the convolution operation; Indicates nonlinear activation; yes , The result of splicing It is the channel scale factor.
2. The BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning according to claim 1, characterized in that: The boundary stream decoder is used to refine the boundaries of objects and output a boundary segmentation map. Specifically, this includes: Use with clear boundaries , , and The feature maps predict lesion boundaries, and three gated convolutional layers are added to disable irrelevant features extracted by the Res encoder; Defined as the first The feature map output by the layer boundary stream decoder, the first Attention map of gated convolutional layers for: ; in, This represents the convolution operation; Indicates nonlinear activation. This represents the channel cascading of the feature map; After passing through a residual block composed of two normalized convolutional layers with skip connections, we obtain Feature map of layer boundary flow decoder : ; in This represents a residual block, which will replace the original... Replace with and The result of the connection; the output of the last layer boundary stream decoder is the boundary segmentation map.
3. The BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning according to claim 1, characterized in that: The segmentation stream decoder generates segmented images by acquiring position and boundary information. Specifically, this includes: the segmented stream decoder using the UNet architecture. , , , Decoding is performed by using dense dilated convolution, pooling, and activation operations to obtain a feature map which is then input to the next layer, ultimately resulting in the segmented image.
4. The BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning according to claim 1, characterized in that, In step one, the preprocessing of the acquired MRI images specifically includes: Preprocessing of MRI images: Rotating MRI images at angles from 30° to 90°, and randomly flipping MRI images horizontally and vertically; z-score normalization of MRI images using mean μ=0 and standard deviation σ=1.
5. The BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning according to claim 1, characterized in that, Step 3: Set loss functions for the position stream decoder, boundary stream decoder, and segmentation stream decoder in the initial segmentation network. Add the loss functions of the position stream decoder, boundary stream decoder, and segmentation stream decoder to obtain the loss function of the initial segmentation network. Specifically, this includes: Loss function of location stream decoder for: ; in, for Corresponding tags; Loss function of boundary stream decoder for: ; in, for Corresponding tags; Loss function of split stream decoder for: ; in, for Corresponding tags; Loss function of the initial segmentation network for: 。 6. The BI-RADS breast lesion segmentation and classification method based on collaborative multi-task learning according to claim 5, characterized in that, When setting loss functions for the classification subnet and the resegmentation subnet, and then summing the loss functions of the classification subnet and the resegmentation subnet to obtain the loss function for the entire segmentation and classification network, the specific steps include: The classification subnetwork uses the original MRI image, along with the obtained location segmentation map, boundary segmentation map, and classification results, and applies binary cross-entropy as the loss function to derive the classification subnetwork's loss function. : ; in Represents the weights of the loss function for each branch of the classification subnet, where , Represents the loss function for each branch of the classification subnet; The reclassified subnet will use binary cross-entropy loss to optimize the final result, resulting in the loss function in the reclassified subnet. : ; M represents the final segmentation result obtained by re-segmenting the subnet; for Corresponding tags; The loss function of the entire segmentation and classification network for: ; , These are the weight parameters.
Citation Information
Patent Citations
Method for detecting X-ray mammary gland lesion image based on feature pyramid network under transfer learning
CN110674866A
MRI multi-modal image segmentation method and system based on ACU-Net
CN112215844A