Ultrasonic image multi-task learning method
By adopting a multi-task learning method with U-type encoder-decoder structure in breast ultrasound image processing, the accuracy challenge of breast ultrasound image multi-task processing in the prior art is solved, and higher segmentation and classification performance and diagnostic consistency are achieved.
Patent Information
- Application Number
- CN202510034584.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art presents accuracy challenges in multitasking of breast ultrasound images and does not fully utilize the synergistic effects of multitasking learning, resulting in information loss and diagnostic uncertainty.
A ultrasonic image multi-task learning method is proposed, and the breast tumor image segmentation and classification model with U-shaped encoder-decoder structure is adopted, including multi-task bottleneck blocks, feature aggregation blocks and ring training methods, which improve the accuracy of segmentation and classification through synergistic effects between features enhancement and tasks.
By effectively exploring the correlation between tasks, the synergy between tumor segmentation and classification is enhanced, the segmentation and classification performance of breast ultrasound images is improved, and the risk of misdiagnosis and misdiagnosis is reduced.
Smart Images

Figure CN119941684A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and in particular relates to a multi-task learning method for ultrasound images. Background Art
[0002] Breast ultrasound is a non-invasive, harmless examination method that can accurately distinguish the nature of breast lumps and detect lesions that are difficult to identify with X-rays in the case of dense breast tissue. However, due to subjective differences among doctors, the diagnosis results may be uncertain. To solve this problem, it is particularly important to introduce automated auxiliary diagnosis methods, which can improve the accuracy and consistency of diagnosis and reduce the risk of misdiagnosis and missed diagnosis.
[0003] Many studies treat segmentation and classification tasks separately without fully considering the relationship between them, which may lead to information loss because segmentation and classification are inherently interrelated. In order to analyze breast ultrasound images more deeply, it is necessary to combine segmentation and classification, which can provide a deeper analysis of breast ultrasound images. Multi-task learning has been widely used in the field of image analysis, especially showing great potential in combining segmentation and classification tasks. However, due to the complexity and difficulty of ultrasound images, existing methods still face challenges in terms of accuracy and do not fully utilize the synergistic effect of multi-task learning.
[0004] In summary, there is an urgent need for a novel and effective multi-task learning method that aims to deepen the interaction between tasks to improve the accuracy of multi-task processing in ultrasound images. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention proposes a multi-task learning method for ultrasound images, which includes: obtaining a breast tumor image and preprocessing it, inputting the preprocessed image into a trained breast tumor image segmentation and classification model for processing, and obtaining a breast tumor image segmentation and classification result;
[0006] The breast tumor image segmentation and classification model is a U-shaped encoder-decoder structure, including a downsampling path, a multi-task bottleneck block, a classification block, a feature aggregation block and an upsampling path; the downsampling path is composed of a first, second, third and fourth encoder, and the upsampling path is composed of a first, second, third and fourth decoder;
[0007] The multi-task bottleneck block connects the downsampling path, classification block and upsampling path, and the downsampling path and the upsampling path are skipped through the feature aggregation block.
[0008] Preferably, the data processing process of the multi-task bottleneck block includes:
[0009] Perform position encoding on the input features to obtain initial classification features and initial segmentation features;
[0010] The self-task module is used to process the initial classification features and the initial segmentation features respectively to obtain the self-attention classification features and the self-attention segmentation features;
[0011] The cross-task module is used to process the initial classification features and initial segmentation features after layer normalization to obtain cross-attention classification features and cross-attention segmentation features;
[0012] The initial classification features, self-attention classification features and cross-attention classification features are added to obtain intermediate classification features; the initial segmentation features, self-attention segmentation features and cross-attention segmentation features are added to obtain intermediate segmentation features;
[0013] The feature enhancement module is used to process the intermediate classification features and the intermediate segmentation features respectively, and the output of the multi-task bottleneck block is obtained, namely the final classification features and segmentation features.
[0014] Furthermore, the self-task module processes the initial classification features and the initial segmentation features respectively as follows:
[0015]
[0016] in, represents the self-attention classification feature, represents the self-attention segmentation feature, represents the initial classification features, represents the initial segmentation features, LN(·) represents layer normalization, and MSA(·) represents multi-head attention.
[0017] Furthermore, the cross-task module processes the initial classification features and the initial segmentation features respectively as follows:
[0018]
[0019] in, represents the cross-attention classification feature, represents the cross-attention segmentation feature, Q cl , K cl and V cl They represent the query, key, and value matrices obtained from the initial classification features after layer normalization; Q seg , K seg and V seg denote the query, key, and value matrices obtained from the initial segmentation features after layer normalization, respectively; B denotes the bias, d denotes the dimension of the query and key matrices, and SoftMax(·) denotes the softmax function.
[0020] Furthermore, the feature enhancement module processes the intermediate classification features and the intermediate segmentation features respectively as follows:
[0021]
[0022] in, represents the classification features output by the multi-task bottleneck block, represents the segmentation features output by the multi-task bottleneck block, represents the mid-level classification features, represents the mid-level segmentation feature, LN(·) represents layer normalization, and MLP(·) represents multi-layer perceptron.
[0023] Preferably, the data processing process of the classification block includes:
[0024] The classification features output by the multi-task bottleneck block are subjected to global average pooling and flattening operations in turn to obtain intermediate features. The intermediate features are processed using a fully connected layer and a sigmoid activation function to obtain the breast tumor image classification results.
[0025] Preferably, the jump connection between the downsampling path and the upsampling path is realized through the feature aggregation block, specifically including: the preprocessed image and the first decoder are connected through the feature aggregation block, the first encoder and the second decoder are connected through the feature aggregation block, the second encoder and the third decoder are connected through the feature aggregation block, and the third encoder and the fourth decoder are connected through the feature aggregation block.
[0026] Furthermore, the data processing process of the feature aggregation block is expressed as:
[0027] f a&m =δ2(l2(δ1(l1(AvgPooling(f e )))+δ1(l1(MaxPooling(f e )))))
[0028]
[0029] in, represents the output features of the feature aggregation block, f a&m represents the activation feature, δ1 and δ2 represent the ReLu and Sigmoid activation functions respectively, l1 represents a fully connected layer with half the size, l2 represents a fully connected layer with incremental size, AvgPooling and MaxPooling represent average pooling and maximum pooling respectively, f e represents the first input of the feature aggregation block, f d Represents the second input to the feature aggregation block.
[0030] Preferably, the total loss function in the breast tumor image segmentation and classification model training process is the weighted sum of the classification loss and the segmentation loss; the segmentation loss is the weighted sum of the cross entropy loss and the dice loss; wherein:
[0031] The dice loss is expressed as:
[0032]
[0033] The classification loss is expressed as:
[0034]
[0035] Among them, L dice (P seg ,y seg ) represents the dice loss, L cl (P cl ,y cl ) represents the classification loss, P seg represents the predicted segmentation map, y seg represents the labeled tumor map, ω represents the weight of malignant cases, P cl Represents the predicted classification probability, y cl represents the true classification label, and γ represents the focusing parameter.
[0036] The beneficial effects of the present invention are:
[0037] The present invention proposes a novel multi-task learning network and introduces a multi-task bottleneck layer. This layer effectively explores the correlation between tasks and enhances the synergy between tumor segmentation and classification, thereby improving the overall performance. The present invention designs a feature aggregation block FAB to optimize the pixel reconstruction process. FAB integrates a channel attention mechanism to ensure the prominence of key features, thereby bridging the semantic gap between the encoder and decoder and improving the accuracy of the segmentation results. The present invention proposes a circular training method that uses the original input and the segmented image predicted by the previous round to help fully capture tumor features, thereby improving the segmentation and classification performance of breast ultrasound images. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Schematic diagram of the structure of the breast tumor image segmentation and classification model in the present invention;
[0039] Figure 2 This is a schematic diagram of the multi-task bottleneck block structure in the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] The present invention proposes a multi-task learning method for ultrasound images, which includes the following contents:
[0042] Breast tumor images are acquired and preprocessed, and the preprocessed images are input into a trained breast tumor image segmentation and classification model for processing to obtain breast tumor image segmentation and classification results.
[0043] The present invention designs a breast tumor image segmentation and classification model on the U-Net network. Figure 1 As shown, the breast tumor image segmentation and classification model is a U-shaped encoder-decoder structure, including a downsampling path, a multi-task bottleneck block, a classification block, a feature aggregation block and an upsampling path; the downsampling path consists of the first, second, third and fourth encoders, and the upsampling path consists of the first, second, third and fourth decoders.
[0044] Preferably, each encoder includes a maximum pooling layer, a convolution layer, a batch normalization layer and a Relu activation function layer. Each decoder includes an upsampling layer, a convolution layer, a batch normalization layer and a Relu activation function layer.
[0045] The multi-task bottleneck block (MTBL) connects the downsampling path, the classification block and the upsampling path; the input of the multi-task bottleneck block is the output of the fourth encoder, and the output of the multi-task bottleneck block is the classification feature and the segmentation feature respectively; the classification feature is used as the input of the classification block, and the classification result of the breast tumor image is obtained after being processed by the classification block; the segmentation feature is used as the input of the fourth decoder, and the segmentation result of the breast tumor image is obtained after being processed by the downsampling path.
[0046] The data processing of the classification block includes:
[0047] The classification features output by the multi-task bottleneck block are subjected to global average pooling and flattening operations in turn to obtain intermediate features. The intermediate features are processed using a fully connected layer and a sigmoid activation function to obtain the breast tumor image classification results.
[0048] like Figure 2 As shown, the data processing process of the multi-task bottleneck block includes:
[0049] Perform position encoding on the input features to obtain the initial classification features and initial segmentation features
[0050] The self-task module is used to classify the initial classification features and initial segmentation features Process and obtain
[0051] Self-attention classification features and self-attention segmentation features The self-task module learns the global context information of each task through task attention. Its processing process is expressed as:
[0052]
[0053] Among them, LN(·) represents layer normalization and MSA(·) represents multi-head attention.
[0054] The cross-task module is used to process the initial classification features and initial segmentation features after layer normalization to obtain the cross-attention classification features. and cross-attention segmentation features The cross-task module captures the dependencies in the input sequence through the attention mechanism, and its processing process is expressed as:
[0055]
[0056] Among them, Q cl , K cl and V cl They represent the query, key, and value matrices obtained from the initial classification features after layer normalization; Q seg , K seg and V seg denote the query, key, and value matrices obtained from the initial segmentation features after layer normalization, respectively; B denotes the bias, d denotes the dimension of the query and key matrices, and SoftMax(·) denotes the softmax function.
[0057] Add the initial classification features, self-attention classification features, and cross-attention classification features to obtain the intermediate classification features Add the initial segmentation feature, self-attention segmentation feature and cross-attention segmentation feature to obtain the intermediate segmentation feature
[0058] The feature enhancement module is used to process the intermediate classification features and intermediate segmentation features respectively to obtain the output of the multi-task bottleneck block, namely the final classification features and segmentation features; the feature enhancement module can realize feature sharing and enhancement between classification and segmentation tasks, thereby significantly improving the performance of each task; the processing process is expressed as:
[0059]
[0060] in, represents the classification features output by the multi-task bottleneck block, represents the segmentation features output by the multi-task bottleneck block, and MLP(·) represents a multi-layer perceptron.
[0061] The downsampling path and the upsampling path are connected by a feature aggregation block (FAB). Specifically, the present invention is implemented by using four feature aggregation blocks with the same structure, wherein the preprocessed image and the first decoder are connected by a first feature aggregation block, the first encoder and the second decoder are connected by a second feature aggregation block, the second encoder and the third decoder are connected by a third feature aggregation block, and the third encoder and the fourth decoder are connected by a fourth feature aggregation block.
[0062] The feature aggregation block channel attention is used to capture the upsampled decoder features and embed them into the features obtained from the encoder to highlight the important and more representative channels in the entire encoder feature map; the output features of each feature aggregation block are used as an input feature of the corresponding decoder.
[0063] The feature aggregation block can be expressed as:
[0064] f a&m =δ2(l2(δ1(l1(AvgPooling(f e )))+δ1(l1(MaxPooling(f e )))))
[0065]
[0066] in, represents the output features of the feature aggregation block, f a&m represents the activation feature, δ1 and δ2 represent the ReLu and Sigmoid activation functions respectively, l1 represents a fully connected layer with half the size, l2 represents a fully connected layer with incremental size, AvgPooling and MaxPooling represent average pooling and maximum pooling respectively, f e represents the first input of the feature aggregation block, i.e., the preprocessed image or the output feature of the encoder, f d Represents the second input of the feature aggregation block, namely the upsampling layer output feature.
[0067] The last f a&m With f d The element-by-element multiplication of allows the encoder features to be weighted and adjusted according to the weights of the decoder features on each channel, and this weighted process helps to strengthen the encoder feature information that is semantically related to the decoder features, improving the richness and discriminability of the feature representation.
[0068] During the training of the breast tumor image segmentation and classification model, a breast ultrasound dataset is prepared, breast tumor labels (benign, malignant) and breast lesion areas are annotated, and corresponding image data preprocessing (such as image normalization and histogram equalization) is performed.
[0069] In order to further utilize the role of tasks, a cyclic training strategy guided by global confidence is introduced to improve the accuracy of segmentation and classification. Ring training is used in the model training process, that is, in the first training, the breast ultrasound image is input into the network, and the network outputs the segmentation probability map. Subsequent training adjusts the input in a weighted manner (based on the segmentation probability map of the previous round and the calculated confidence), thereby helping the network to focus more accurately on difficult-to-segment areas and improve the performance of the model. Specifically, a confidence function is designed to represent the confidence level of the pixel point. The confidence calculation formula is:
[0070] c i = abs(y i -0.5)*2
[0071] Among them, c i is the confidence of the pixel, y i is the predicted probability value, and abs represents the absolute value function. When the predicted value is close to 0 or 1, the confidence is close to 1, indicating that the model is very confident in the classification of the pixel; when the predicted value is close to 0.5, the confidence is close to 0, indicating that the model has low confidence in the classification of the pixel. Based on the confidence of all pixels, the global confidence c of the image is calculated by the following formula:
[0072]
[0073] Multiply the segmentation probability map of the previous round by the corresponding confidence value, and then add the result to the original image as the input of the current round.
[0074] input=img+SegProbMap i-1 *c
[0075] Among them, input represents input, img represents the original image, SegProbMap i-1 Represents the segmentation probability map of the previous round.
[0076] The loss function in the training process of the breast tumor image segmentation and classification model designed by the present invention is the weighted sum of the classification loss and the segmentation loss:
[0077] L joint =ηL cl +(1-η)L seg
[0078] Among them, Ljont represents the multi-task loss, i.e. the total loss of the model, L cl represents the classification loss, L seg represents the segmentation loss, and η represents the weight of the classification loss.
[0079] Segmentation loss:
[0080] For segmentation tasks, Dice loss has a better handle on class imbalance. Since it directly measures the overlap between the predicted area and the true area, it can also be effectively optimized for small objects. The cross entropy loss takes into account the classification results of each pixel and can globally optimize the segmentation task. The hybrid loss function that combines Dice loss and cross entropy loss can make full use of the advantages of both to improve the performance of the image segmentation model. The segmentation loss is defined as the cross entropy loss L ce With dice loss L dice The weighted sum of:
[0081] L seg =βL ce +(1-β)L dice
[0082] Among them, β represents the weight of cross entropy loss.
[0083] The cross entropy loss is expressed as:
[0084]
[0085] Where N is the total number of voxels, y i represents the true voxel value of the i-th pixel, y i ′ represents the i-th predicted voxel value.
[0086] The dice loss is expressed as:
[0087]
[0088] Among them, P seg represents the predicted segmentation map, y seg Represents a labeled tumor map.
[0089] Classification loss:
[0090] Generally, unbalanced datasets are one of the main problems in classification. This imbalance may cause the model to over-focus on the majority class and ignore the minority class during training. In order to deal with this imbalance, the present invention adopts focal loss as the classification loss function, which is expressed as:
[0091]
[0092] Among them, ′ represents the weight of malignant cases, P clRepresents the predicted classification probability, y cl represents the true classification label, and γ represents the focusing parameter.
[0093] After training, the optimal breast tumor image segmentation and classification model is obtained. By inputting the ultrasound image into the optimal generation model, the tumor classification and lesion area segmentation results can be obtained.
[0094] In summary, the present invention proposes a novel multi-task learning network and introduces a multi-task bottleneck layer. This layer effectively explores the correlation between tasks and enhances the synergy between tumor segmentation and classification, thereby improving the overall performance. During the segmentation process, there are semantic differences between the encoder and decoder subnetworks, which affects the accuracy of reconstruction. To solve this problem, the present invention designs a feature aggregation block FAB to optimize the pixel reconstruction process. FAB integrates a channel attention mechanism to ensure the prominence of key features, thereby bridging the semantic gap between the encoder and decoder and improving the accuracy of the segmentation results. The complexity of tissue types and the diversity of lesion morphology in breast ultrasound images make it difficult to accurately extract tumor features. To address this challenge, the present invention proposes a cyclic training method that uses the original input and the segmented image predicted by the previous round to help comprehensively capture tumor features, thereby improving the segmentation and classification performance of breast ultrasound images.
[0095] Evaluation of the present invention:
[0096] For the segmentation task. The proposed method is compared with seven representative methods that are currently the most advanced in this field, namely Unet (Ronneberger et al., 2015), DCSAU-Net (Xu, Ma, Na and Duan, 2023), TransUnet (Chen et al., 2021), Unet++ (Zhou et al., 2018), UCTransNet (Wang et al., 2022), Nu Net (Chen et al., 2023) and ESKnet (Chen et al., 2024). These methods have shown excellent performance in the field of breast tumor segmentation. Table 1 lists the evaluation results of these segmentation methods on the BUSI dataset, where the bold values represent the best results and the values in the lower right corner represent the degree of improvement. It can be seen from the evaluation indicators listed in the table that the proposed method outperforms other algorithms on the BUSI dataset. Specifically, it outperforms the runner-up method by 0.39%, 0.16% and 0.45 in DSC (Dice similarity coefficient), JI (intersection over union) and HD, respectively.
[0097] Table 1 Segmentation comparison results on the BUSI dataset
[0098] Methods DSC JI HD Unet <![CDATA[77.81 ↑4.14 ]]> <![CDATA[68.87 ↑4.44 ]]> <![CDATA[35.94 ↓2.21 ]]> DCSAU-Net <![CDATA[78.94 ↑3.01 ]]> <![CDATA[69.79 ↑3.52 ]]> <![CDATA[28.99 ↓5.26 ]]> Unet++ <![CDATA[78.94 ↑3.01 ]]> <![CDATA[70.28 ↑3.03 ]]> <![CDATA[32.53 ↓8.80 ]]> TransUnet <![CDATA[79.37 ↑2.58 ]]> <![CDATA[71.15 ↑2.16 ]]> <![CDATA[27.32 ↓3.59 ]]> UCTransNet <![CDATA[80.971 ↑0.98 ]]> <![CDATA[72.25 ↑1.06 ]]> <![CDATA[27.74 ↓3.01 ]]> Nu_Net <![CDATA[81.09 ↑0.86 ]]> <![CDATA[72.83 ↑0.48 ]]> <![CDATA[27.77 ↓4.04 ]]> ESKnet <![CDATA[81.56 ↑0.39 ]]> <![CDATA[73.15 ↑0.16 ]]> <![CDATA[24.18 ↓0.45 ]]> ours 81.95 73.31 23.73
[0099] In order to verify the performance of the method of the present invention in classification tasks, especially the performance of breast cancer medical image classification, we conducted a comprehensive comparative analysis and compared the method of the present invention with seven widely recognized and leading methods in the field, including Resnet50 (He et al., 2016), Efficient (Tan and Le, 2019), Regnet (Radosavovic et al., 2020), Vit (Dosovitskiy et al., 2020), Swin-transformer (Liu et al., 2021), MaxVit (Tu et al., 2022), Medvit (Manzari et al., 2023). In the evaluation of Table 2, the method of the present invention showed excellent performance on the BUSI dataset, and achieved a leading position in the three key indicators of ACC (accuracy), F1 and AUC, which were 88.09%, 82.13% and 94.31% respectively. The method of the present invention is significantly better than the second-place MaxVit method, which is improved by 2% in ACC and 2.9% in F1 score.
[0100] Table 2 Classification comparison results on the BUSI dataset
[0101] Methods F1 ACC AUC Vit <![CDATA[62.80 ↑19.33 ]]> <![CDATA[74.34 ↑13.75 ]]> <![CDATA[79.09 ↑15.22 ]]> Swin-transformer <![CDATA[69.56 ↑12.57 ]]> <![CDATA[79.44 ↑8.65 ]]> <![CDATA[87.99 ↑6.32 ]]> Efficient <![CDATA[72.20 ↑9.93 ]]> <![CDATA[82.68 ↑5.40 ]]> <![CDATA[89.72 ↑4.59 ]]> RegNet <![CDATA[75.28 ↑6.85 ]]> <![CDATA[83.45 ↑4.64 ]]> <![CDATA[91.49 ↑2.82 ]]> MedVit <![CDATA[75.85 ↑6.28 ]]> <![CDATA[84.39 ↑3.70 ]]> <![CDATA[91.27 ↑3.04 ]]> RestNet50 <![CDATA[78.64 ↑3.49 ]]> <![CDATA[85.16 ↑2.93 ]]> <![CDATA[93.38 ↑0.93 <!-- 7 -->]]> MaxVit <![CDATA[79.23 ↑2.90 ]]> <![CDATA[86.09 ↑2.00 ]]> <![CDATA[91.65 ↑2.66 ]]> ours 82.13 88.09 94.31
[0102] The above embodiments further illustrate the purpose, technical solutions and advantages of the present invention in detail. It should be understood that the above embodiments are only preferred implementation modes of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made to the present invention within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-task learning method for ultrasound images, characterized in that: include: Acquire breast tumor images and preprocess them, input the preprocessed images into a trained breast tumor image segmentation and classification model for processing, and obtain breast tumor image segmentation and classification results; The breast tumor image segmentation and classification model is a U-shaped encoder-decoder structure, including a downsampling path, a multi-task bottleneck block, a classification block, a feature aggregation block and an upsampling path; the downsampling path is composed of a first, second, third and fourth encoder, and the upsampling path is composed of a first, second, third and fourth decoder; The multi-task bottleneck block connects the downsampling path, classification block and upsampling path, and the downsampling path and the upsampling path are skipped through the feature aggregation block.
2. The multi-task learning method for ultrasound images according to claim 1, characterized in that: The data processing process of the multi-task bottleneck block includes: Perform position encoding on the input features to obtain initial classification features and initial segmentation features; The self-task module is used to process the initial classification features and the initial segmentation features respectively to obtain the self-attention classification features and the self-attention segmentation features; The cross-task module is used to process the initial classification features and initial segmentation features after layer normalization to obtain cross-attention classification features and cross-attention segmentation features; The initial classification features, self-attention classification features and cross-attention classification features are added to obtain intermediate classification features; the initial segmentation features, self-attention segmentation features and cross-attention segmentation features are added to obtain intermediate segmentation features; The feature enhancement module is used to process the intermediate classification features and the intermediate segmentation features respectively, and the output of the multi-task bottleneck block is obtained, namely the final classification features and segmentation features.
3. The ultrasonic image multi-task learning method according to claim 2, characterized in that: The self-task module processes the initial classification features and initial segmentation features respectively as follows: in, represents the self-attention classification feature, represents the self-attention segmentation feature, represents the initial classification features, represents the initial segmentation features, LN(·) represents layer normalization, and MSA(·) represents multi-head attention.
4. The ultrasonic image multi-task learning method according to claim 2, characterized in that: The cross-task module processes the initial classification features and initial segmentation features respectively as follows: in, represents the cross-attention classification feature, represents the cross-attention segmentation feature, Q cl , K cl and V cl They represent the query, key, and value matrices obtained from the initial classification features after layer normalization; Q seg , K seg and V seg denote the query, key, and value matrices obtained from the initial segmentation features after layer normalization, respectively; B denotes the bias, d denotes the dimension of the query and key matrices, and SoftMax(·) denotes the softmax function.
5. The ultrasonic image multi-task learning method according to claim 2, characterized in that: The feature enhancement module processes the intermediate classification features and the intermediate segmentation features respectively as follows: in, represents the classification features output by the multi-task bottleneck block, represents the segmentation features output by the multi-task bottleneck block, represents the mid-level classification features, represents the mid-level segmentation feature, LN(·) represents layer normalization, and MLP(·) represents multi-layer perceptron.
6. The ultrasound image multi-task learning method according to claim 1, characterized in that: The data processing process of the classification block includes: The classification features output by the multi-task bottleneck block are subjected to global average pooling and flattening operations in turn to obtain intermediate features. The intermediate features are processed using a fully connected layer and a sigmoid activation function to obtain the breast tumor image classification results.
7. The ultrasound image multi-task learning method according to claim 1, characterized in that: The jump connection between the downsampling path and the upsampling path is realized through the feature aggregation block, which specifically includes: the preprocessed image and the first decoder are connected by a feature aggregation block, the first encoder and the second decoder are connected by a feature aggregation block, the second encoder and the third decoder are connected by a feature aggregation block, and the third encoder and the fourth decoder are connected by a feature aggregation block.
8. The ultrasound image multi-task learning method according to claim 7, characterized in that: The data processing process of the feature aggregation block is expressed as: f a&m =δ2(l2(δ1(l1(AvgPooling(f e )))+δ1(l1(MaxPooling(f e ))))) in, represents the output features of the feature aggregation block, f a&m represents the activation feature, δ1 and δ2 represent the ReLu and Sigmoid activation functions respectively, l1 represents a fully connected layer with half the size, l2 represents a fully connected layer with incremental size, AvgPooling and MaxPooling represent average pooling and maximum pooling respectively, f e represents the first input of the feature aggregation block, f d Represents the second input to the feature aggregation block.
9. The ultrasound image multi-task learning method according to claim 1, characterized in that: The total loss function in the training process of breast tumor image segmentation and classification model is the weighted sum of classification loss and segmentation loss; segmentation loss is the weighted sum of cross entropy loss and dice loss; where: The dice loss is expressed as: The classification loss is expressed as: Among them, L dice (P seg ,y seg ) represents the dice loss, L cl (P cl ,y cl ) represents the classification loss, P seg represents the predicted segmentation map, y seg represents the labeled tumor map, ω represents the weight of malignant cases, P cl Represents the predicted classification probability, y cl represents the true classification label, and γ represents the focusing parameter.
Citation Information
Cited By
Breast tumor image multi-task learning method, system, device and medium
CN121436076A
A multi-task learning method, system, device and medium for breast tumor images
CN121436076B