Multitask model training method, image processing method, electronic equipment and storage medium
Through the multi-task model training method, the segmentation and classification loss are calculated using the shared encoder and self-attention mechanism, and the accuracy of ultrasound images is solved when detecting gallbladder polyps, achieving efficient classification of gallbladder polyps.
Patent Information
- Application Number
- CN202510406186.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
Ultrasound images have poor image quality, noise interference, and resolution limitations when detecting gallbladder polyps, which leads to doctors relying on experience and intuition during diagnosis and inaccurate classification.
The multi-task model training method is adopted, and the cropped image samples are input into the segmentation network and classification network through a shared encoder, segmentation loss and classification loss are calculated, model parameters are updated using comprehensive loss, and combined with the self-attention mechanism and multi-scale feature fusion to improve the classification accuracy of the model.
The accuracy of gallbladder polyps classification is improved, and through mutual guidance training between segmentation loss and classification loss, the virtuous cycle improvement of classification and segmentation tasks is promoted, and the utilization of features between different scales is enhanced.
Smart Images

Figure CN120339691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a multi-task model training method, an image processing method, an electronic device, and a storage medium. Background Art
[0002] Gallbladder polyps refer to tissues protruding from the inner wall of the gallbladder, usually detected in ultrasound examinations, and are divided into benign polyps (such as cholesterol polyps) and malignant or suspected malignant polyps (such as adenomas, cancers). Accurately distinguishing true and false polyps (i.e., benign polyps from malignant or cancerous polyps) has important clinical significance for treatment decisions, surgical plan formulation, and patient prognosis.
[0003] However, ultrasound images often have problems such as poor image quality, noise interference, and resolution limitations when detecting gallbladder polyps, resulting in doctors relying on experience and intuition in diagnosis and inaccurate classification of gallbladder polyps. Summary of the Invention
[0004] In view of this, the purpose of the embodiments of the present invention is to provide a multi-task model training method, an image processing method, an electronic device, and a storage medium to at least partially improve the above problems.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:
[0006] In a first aspect, the embodiments of the present invention provide a multi-task model training method. The multi-task model includes a shared encoder, a segmentation network, and a classification network. The method includes:
[0007] Input the cropped image sample into the shared encoder to obtain a first feature map; the cropped image sample includes an annotated classification label and an annotated segmentation label;
[0008] Input the first feature map into the segmentation network to obtain a segmentation result, and calculate a segmentation loss according to the segmentation result and the annotated segmentation label;
[0009] Input the first feature map into the classification network to obtain a classification result, and calculate a classification loss according to the classification result and the annotated classification label;
[0010] Calculate a comprehensive loss of the multi-task model according to the classification loss and the segmentation loss, and update the parameters of the multi-task model using the comprehensive loss.
[0011] Optionally, the classification network includes at least one regressor and a classifier; the step of inputting the first feature map into the classification network to obtain a classification result and calculating a classification loss according to the classification result and the annotated classification label includes:
[0012] Input the first feature map into each of the regressors respectively, and encode them into first vectors respectively;
[0013] Concatenate the first vectors by dimension to obtain a second vector, and input the second vector into the classifier to obtain the classification result of the cropped image sample;
[0014] Based on the classification loss function, calculate the classification loss according to the classification result and the labeled classification label.
[0015] Optionally, the classification network includes at least one regressor, a fully connected layer, and a classifier; the step of inputting the first feature map into the classification network to obtain a classification result and calculating the classification loss according to the classification result and the labeled classification label includes:
[0016] Extract the morphological features of the segmentation result; the morphological features include morphological sub-features with the same number as the number of regressors;
[0017] Input the first feature map into each of the regressors respectively, and encode them into first vectors respectively;
[0018] Input each of the first vectors into the fully connected layer respectively to obtain the predicted values of each of the first vectors;
[0019] Based on the regression loss function, calculate the regression loss according to the predicted values and the morphological sub-features;
[0020] Concatenate the first vectors by dimension to obtain a second vector, and input the second vector into the classifier to obtain the classification result of the cropped image sample;
[0021] Based on the classification loss function, calculate the classification loss according to the classification result, the labeled classification label, and the regression loss.
[0022] Optionally, the regression loss function is:
[0023]
[0024] where y i is the morphological sub-feature, is the predicted value, δ is a hyperparameter, L Huber is the sub-regression loss of a single regressor, L Res is the regression loss, λ j is the weight of the j-th sub-regression loss, and n is the number of regressors.
[0025] Optionally, the shared encoder includes at least three encoding blocks, the first feature map includes at least two second feature maps and a high-order feature map, the high-order feature map is the feature map output by the last encoding block, the segmentation network includes a multi-scale feature fusion self-attention module and a decoder, and the decoder includes decoding blocks with the same number as each of the encoding blocks; inputting the first feature map into the segmentation network to obtain a segmentation result, and calculating a segmentation loss according to the segmentation result and the labeled segmentation label, includes:
[0026] Input each of the second feature maps into the multi-scale feature fusion self-attention module to obtain each third feature map encoded by self-attention;
[0027] Input each of the third feature maps and the high-order feature map into the corresponding decoding block of the decoder to obtain the segmentation result of the cropped image sample;
[0028] Based on a segmentation loss function, calculate a segmentation loss according to the segmentation result and the labeled segmentation label.
[0029] Optionally, the inputting each of the second feature maps into the multi-scale feature fusion self-attention module to obtain each third feature map encoded by self-attention includes:
[0030] Fuse each of the second feature maps to obtain a fused feature map;
[0031] Divide the fused feature map into multiple window sub-features;
[0032] Based on a position embedding formula, perform position embedding on each of the window sub-features to obtain position-embedded features after position embedding;
[0033] Perform self-attention attention encoding on the position-embedded features to obtain a self-attention feature map;
[0034] Restore the position of the self-attention feature map to each third feature map corresponding to each of the second feature maps.
[0035] Optionally, the position embedding formula is:
[0036]
[0037] where z0 is the position-embedded feature, is the k-th window sub-feature, E is a hidden vector, E pos is a position encoding matrix.
[0038] In a second aspect, an embodiment of the present invention provides an image processing method, the method includes:
[0039] Obtain the cropped image to be processed;
[0040] Input the cropped image to be processed into the multi-task model to obtain the segmentation result and classification result of the cropped image to be processed; the multi-task model is trained by the multi-task model training method described in any item of the first aspect.
[0041] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, the method described in any of the above items is implemented.
[0042] In a fourth aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any of the above items is implemented.
[0043] A multi-task model training method, an image processing method, an electronic device, and a storage medium provided by an embodiment of the present invention calculate their respective losses through a shared encoder of a segmentation network and a classification network, and are trained according to the mutual guidance of the segmentation loss and the classification loss, thereby improving the classification accuracy of the model.
[0044] Furthermore, the morphological features extracted from the output result of the segmentation model are used as pseudo ground truths to optimize the regression model. The pseudo ground truths become more accurate as the segmentation result is continuously optimized, promoting a virtuous cycle of improvement between the classification task and the segmentation task. Through the multi-scale feature fusion self-attention mechanism, the multi-scale gallbladder polyp ultrasound image feature maps are extracted from the encoder for fusion, and the self-attention mechanism is used to model the long-range characteristics of the fusion features, enhancing the mutual utilization degree between different scale features.
[0045] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic structural block diagram of an electronic device provided by an embodiment of the present invention;
[0048] Figure 2 It is a schematic structural block diagram of a multi-task model provided by an embodiment of the present invention;
[0049] Figure 3 Schematic flowchart of a multi-task model training method provided by an embodiment of the present invention;
[0050] Figure 4 Another schematic structural block diagram of a multi-task model provided by an embodiment of the present invention;
[0051] Figure 5 Another schematic flowchart of a multi-task model training method provided by an embodiment of the present invention;
[0052] Figure 6 Schematic flowchart of step S221 provided by an embodiment of the present invention;
[0053] Figure 7 Schematic diagram of feature fusion provided by an embodiment of the present invention;
[0054] Figure 8 Schematic diagram of position restoration provided by an embodiment of the present invention;
[0055] Figure 9 Another schematic structural block diagram of a multi-task model provided by an embodiment of the present invention;
[0056] Figure 10 Schematic flowchart of step S230 provided by an embodiment of the present invention;
[0057] Figure 11 Another schematic structural block diagram of a multi-task model provided by an embodiment of the present invention;
[0058] Figure 12 Another schematic flowchart of step S230 provided by an embodiment of the present invention;
[0059] Figure 13 Schematic diagram of morphological features provided by an embodiment of the present invention;
[0060] Figure 14 Flowchart of a multi-task model training method provided by an embodiment of the present invention;
[0061] Figure 15 Schematic flowchart of an image processing method provided by an embodiment of the present invention.
[0062] Icons: 100 - Electronic device; 101 - Memory; 102 - Communication interface; 103 - Processor; 104 - Communication bus; 400 - Multitasking model; 410 - Shared encoder; 411 - Encoding block; 420 - Classification network; 421 - Regressor; 422 - Classifier; 423 - Fully connected layer; 430 - Segmentation network; 431 - Multi-scale feature fusion self-attention module; 432 - Decoder; 4321 - Decoding block. Detailed implementation
[0063] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0064] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but is merely representative of selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of the present invention.
[0065] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings. At the same time, in the description of the present invention, terms such as "first", "second", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.
[0066] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0067] As described in the background art, it is crucial to classify the gallbladder polyps found in ultrasonic examinations as benign or malignant. However, ultrasonic images often have problems such as poor image quality, noise interference, and resolution limitations when detecting gallbladder polyps, resulting in doctors relying on experience and intuition during diagnosis and inaccurate classification of gallbladder polyps.
[0068] Based on the above situation, the embodiments of the present invention provide a multi-task model training method, an image processing method, an electronic device, and a storage medium. By using a shared encoder for the segmentation network and the classification network, their respective losses are calculated, and training is carried out based on the mutual guidance of the segmentation loss and the classification loss, thereby improving the classification accuracy of the model.
[0069] To implement the process steps and functions of each example of the present invention, please refer to Figure 1 , Figure 1 which is a schematic structural block diagram of an electronic device provided by an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103, and the memory 101 and the processor 103 are electrically connected directly or indirectly to each other to achieve data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101 to perform various functional applications and data processing.
[0070] The electronic device 100 can be, but is not limited to, a personal computer (PC), a server, a distributed computer, etc. It can be understood that the electronic device 100 is not limited to a physical server, but can also be a virtual machine on a physical server, a virtual machine built on a cloud platform, etc., which can provide the same functions as the server or virtual machine. The operating system of the electronic device 100 can be, but is not limited to, the Windows system, the Linux system, etc.
[0071] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.
[0072] The communication connection between the electronic device 100 and an external device is implemented through at least one communication interface 102 (which can be wired or wireless).
[0073] The processor 103 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the embodiments of the present invention may be completed by the integrated logic circuit in the hardware of the processor 103 or instructions in the form of software. The processor 103 may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0074] It can be understood that Figure 1 The structure shown is only schematic, and the electronic device 100 may also include more or fewer components than those Figure 1 shown, or have a configuration different from that Figure 1 shown. Figure 1 Each component shown can be implemented by hardware, software, or a combination thereof.
[0075] Next, an exemplary description is provided for the multi-task model training method provided by the present invention. The execution subject of this method may be the above-mentioned Figure 1 shown electronic device 100. Referring to Figure 2 , the multi-task model 400 includes a shared encoder 410, a segmentation network 430, and a classification network 420. Referring to Figure 3 , this method includes the following steps as Figure 3 described:
[0076] S210: Input the cropped image sample into the shared encoder to obtain a first feature map; the cropped image sample includes an annotated classification label and an annotated segmentation label.
[0077] S220: Input the first feature map into the segmentation network to obtain a segmentation result, and calculate a segmentation loss based on the segmentation result and the annotated segmentation label.
[0078] S230: Input the first feature map into the classification network to obtain a classification result, and calculate a classification loss based on the classification result and the annotated classification label.
[0079] S240: Calculate the comprehensive loss of the multi-task model based on the classification loss and the segmentation loss, and update the parameters of the multi-task model using the comprehensive loss.
[0080] In this embodiment, the ultrasonic image of gallbladder polyps is used for exemplary illustration. The cropped image sample of gallbladder polyps is the polyp image after cropping, and this cropped image has richer prior feature information of the image (such as the boundary, color, etc. of gallbladder polyps). This cropped image sample can be the result detected from the gallbladder ultrasonic image by a target detection model (such as YOLO). Then, label the cropped image sample, label the benign and malignant classification of the cropped image sample and mark the specific position of the gallbladder polyp. For example, the classification of a cropped image is benign, and the mark for the polyp is a box, with the pixel value inside the box being 1 and outside the box being 0.
[0081] In this multi-task model, the classification network and the segmentation network share an encoder, such as Figure 2 as shown. Input the cropped image sample into this shared encoder, obtain the first feature map after encoding, then input this first feature map into the segmentation network and the classification network respectively, obtain the segmentation result and the classification result, calculate their respective segmentation losses and classification losses, and calculate the comprehensive loss based on the segmentation loss and the classification loss, and perform iterative update on this multi-task model according to the comprehensive loss.
[0082] This method calculates their respective losses through the shared encoder of the segmentation network and the classification network, and conducts training according to the mutual guidance of the segmentation loss and the classification loss, so as to improve the classification accuracy of the model.
[0083] For the segmentation of the first feature map in step S220, in order to obtain richer information in the first feature map, the self-attention mechanism can be used to focus on the features at different stages and levels during the encoding process of the gallbladder polyp image for long-range attention. Therefore, in a possible implementation manner, referring to Figure 4 , the shared encoder 410 can include at least three encoding blocks 411, the first feature map includes at least two second feature maps and a high-order feature map, the high-order feature map is the feature map output by the last encoding block, the segmentation network 430 includes a multi-scale feature fusion self-attention module 431 and a decoder 432, and the decoder includes decoding blocks 4321 with the same number as each encoding block. Referring to Figure 5 , step S220 can include the following steps:
[0084] S221: Input each second feature map into the multi-scale feature fusion self-attention module to obtain each third feature map encoded by self-attention.
[0085] S222: Input each third feature map and the high-order feature map into the corresponding decoding block of the decoder to obtain the segmentation result of the cropped image sample.
[0086] S223: Calculate the segmentation loss based on the segmentation loss function, the segmentation result, and the labeled segmentation label.
[0087] The shared encoder has multiple encoding blocks. The cropped image is input into the first encoding block to obtain a second feature map, and then the second feature map is input into the next encoding block to obtain another second feature map. This process continues until the feature map is input into the last encoding block. The finally output feature map is a high-order feature map. Each second feature map is input into the multi-scale feature fusion self-attention module. After self-attention encoding of each second feature map, the corresponding third feature map is obtained. Each third feature map and the high-order feature map are input into the corresponding decoding block of the decoder. After being processed by the decoder, the segmentation result of the cropped image sample is obtained through the segmentation head. Finally, based on the segmentation loss function, the segmentation loss can be calculated according to the segmentation result and the labeled segmentation label.
[0088] The segmentation loss function can be composed of the Dice loss function and the cross-entropy loss function. The Dice loss can effectively alleviate the negative impact of foreground imbalance in the samples on the model learning, and pays more attention to the mining of the target. However, there will be a problem of loss saturation. Therefore, the cross-entropy loss is used in combination to balance this problem. Optionally, the segmentation loss function can be expressed as follows:
[0089] L Seg = L Dice + L BCE
[0090] where L Seg is the segmentation loss, L Dice is the Dice loss, and L BCE is the cross-entropy loss.
[0091] The Dice loss is:
[0092]
[0093] where y i is the ground truth of each pixel point in the cropped image, with a value of 1 within the labeled segmentation label and 0 outside the labeled segmentation label. is the value of each pixel point in the predicted segmentation result, with a value of 1 within the segmentation box and 0 outside the segmentation box.
[0094] The cross-entropy loss is:
[0095]
[0096] where N is the number of pixel points in the cropped image, y i is the ground truth of each pixel point in the cropped image, and p iis the predicted probability that each pixel in the cropped image is the true value. For example, if the predicted probability that the i-th pixel is the true value is 60%, then p i is 0.6, and since 60% is greater than 50%, then is 1. If the predicted probability that the i-th pixel is the true value is 40%, then p i is 0.4, and since 40% is less than 50%, then is 0.
[0097] There are various ways to perform self-attention encoding in step S221. For example, the self-attention mechanism of the Transformer model. Optionally, refer to Figure 6 , step S221 may include the following steps:
[0098] S2211: Fuse each second feature map to obtain a fused feature map.
[0099] S2212: Divide the fused feature map into multiple window sub-features.
[0100] S2213: Based on the position embedding formula, perform position embedding on each window sub-feature to obtain the position-embedded feature after position embedding.
[0101] S2214: Perform self-attention encoding on the position-embedded feature to obtain a self-attention feature map.
[0102] S2215: Restore the position of the self-attention feature map to each third feature map corresponding to each second feature map.
[0103] The embodiment of the present invention is based on the traditional Unet structure and is improved. Using a self-attention mechanism for multi-scale feature fusion, the encoded features at different stages of the cropped image are fused, and the self-attention mechanism is used to focus on the features at different stages and levels during the encoding process of the gallbladder polyp image for long-range attention. Exemplarily, if there are 5 encoding blocks in the shared encoder, the output result of the 5th layer encoding block is a high-order feature map. First, fuse the 4 feature maps of different scales output by the first 4 encoding blocks. The fusion process can be referred to Figure 7 , input the output result of the first layer encoding block into the second layer encoding block for encoding, and perform average pooling and convolution processing on the output result of the first layer encoding block. Then add the obtained result to the output result of the second layer encoding block, and perform average pooling and convolution processing on the added result. Add the obtained result to the output result of the third layer encoding block, and perform average pooling and convolution processing on the added result. Add the obtained result to the output result of the fourth layer encoding block, and perform two 1×1 convolution block processing on the added result to finally obtain a fused feature map.
[0104] After obtaining the fused feature map, perform position embedding on it. First, divide the fused feature map into multiple window sub-features. The sizes of the window sub-features can be the same, for example, 3×3, that is, the window sub-feature is a nine-square grid including 9 pixel points. The sizes of the window sub-features can also be different, for example, some are 2×2, some are 3×3, some are 5×5, etc.
[0105] To encode the spatial information of the window sub-features, specific position embedding needs to be introduced to retain the position information. Optionally, position embedding can be performed by the following position embedding formula:
[0106]
[0107] where \(z_0\) is the position embedding feature, is the \(k\)-th window sub-feature, \(E\) is the hidden vector, and \(E\) pos is the position encoding matrix.
[0108] By performing self-attention encoding on the position embedding feature through the self-attention mechanism, the self-attention feature map can be obtained. Finally, the self-attention feature map is positionally restored to the corresponding third feature maps of each second feature map. Exemplarily, the position restoration process can be referred to Figure 8 , first input the self-attention feature map into a 3×3 convolutional block to obtain the feature that needs to be input into the fourth decoding block later. Then, after processing this feature through a 2×2 deconvolutional block and a 3×3 convolutional block, the feature that needs to be input into the third decoding block later can be obtained. Continuing to process it through a 2×2 deconvolutional block and a 3×3 convolutional block in sequence, the corresponding third feature maps of each second feature map can be obtained.
[0109] The structure of the classification network can be various. In one possible implementation, for example, Figure 9 , the classification network 420 includes at least one regressor 421 and a classifier 422. Refer to Figure 10 , this step S230 may include the following steps:
[0110] S231: Input the first feature map into each regressor respectively and encode them into first vectors.
[0111] S232: Concatenate the first vectors by dimension to obtain a second vector, and input the second vector into the classifier to obtain the classification result of the cropped image sample.
[0112] S233: Based on the classification loss function, calculate the classification loss according to the classification result and the labeled classification label.
[0113] Exemplarily, if the number of regressors is 7, the first feature map is input into the 7 regressors respectively, encoded into 7 different 1×256-dimensional first vectors, the 7 first vectors are concatenated into a 1×1792-dimensional second vector in this dimension, and then the second vector is input into the classifier to obtain the classification result of the cropped image sample, that is, the gallbladder polyp here is benign or malignant. The classifier can be a Deep Neural Network (DNN). Finally, the classification loss is calculated according to the classification result and the labeled classification label. Since the distribution of benign and malignant gallbladder polyps is unbalanced at the data level, to deal with the data imbalance problem and distinguish easy and difficult samples, the classification loss function can be as follows:
[0114] L Focal = -α(1 - p) γ ylog(p) - (1 - α)p γ (1 - y)log(1 - p)
[0115] where L Focal is the classification loss, α is the weight, p is the classification result, and y is the labeled classification label.
[0116] In another possible implementation, to guide the classification network, the segmentation result obtained by the segmentation network can be used as prior information to guide the classification network for classification. Exemplarily, see Figure 11 , the classification network 420 includes at least one regressor 421, a fully connected layer 423, and a classifier 422. See Figure 12 , this step S230 may include the following steps:
[0117] S231`: Extract the morphological features of the segmentation result. Among them, the morphological features include morphological sub-features with the same number as the number of regressors.
[0118] S232`: Input the first feature map into each regressor respectively, and encode them into first vectors.
[0119] S233`: Input each first vector into the fully connected layer respectively to obtain the predicted values of each first vector.
[0120] S234`: Based on the regression loss function, calculate the regression loss according to each predicted value and each morphological sub-feature.
[0121] S235`: Concatenate each first vector by dimension to obtain a second vector, and input the second vector into the classifier to obtain the classification result of the cropped image sample.
[0122] S236`: Based on the classification loss function, calculate the classification loss according to the classification result, the labeled classification label, and the regression loss.
[0123] For the segmentation network to obtain the segmentation result, first extract the morphological features of the segmentation result. The morphological features include multiple morphological sub-features, and the number thereof is the same as the number of regressors. For example, if the number of regressors is 7, there are 7 morphological sub-features. Refer to Figure 13 , the morphological features are as Figure 13 shown. Each morphological sub-feature is respectively the lesion area, the area of the smallest convex polygon corresponding to the lesion, the major axis of the ellipse fitted to the lesion area, the minor axis of the ellipse fitted to the lesion area, the number of contour pixels (perimeter) of the lesion, the angle (direction) between the major axis of the fitted ellipse and the X-axis, and the diameter of the circle equal to the lesion area (equivalent diameter). These 7 morphological features are introduced into the classification module as prior information to guide the model to perform effective classification according to the polyp features.
[0124] Step S232` is the same as the above-mentioned step S231. The first feature map is respectively input into each regressor and encoded into a first vector. If the number of regressors is 7, the first feature map is respectively input into 7 regressors and encoded into 7 different first vectors. After obtaining the first vectors, in one branch, like S232 above, after concatenating the first vectors, input them into the classifier to obtain the classification result. In the other branch, each first vector is respectively input into the fully connected layer to respectively obtain the predicted values of the first vectors. Based on the regression loss function, the regression loss is calculated according to the predicted values and the morphological sub-features.
[0125] Optionally, the regression loss function can be:
[0126]
[0127] where y i is the morphological sub-feature, is the predicted value, δ is the hyperparameter, L Huber is the sub-regression loss of a single regressor, L Res is the regression loss, λ j is the weight of the j-th sub-regression loss, and n is the number of regressors.
[0128] Since occasional outliers may occur when extracting the true values of a large number of morphological feature values, and the morphological states of gallbladder polyps are also numerous, which may also lead to outliers, this loss function can effectively reduce the influence of outliers on the model fitting, and the hyperparameter δ is used to balance the loss of large errors.
[0129] After adding the regression loss, the classification loss can be calculated according to the classification result, the labeled classification label, and the regression loss. The classification loss can be calculated by the following formula:
[0130] L Cls = L Res + L focal
[0131]
[0132] L Focal =-α(1 - p) γ ylog(p)-(1 - α)p γ (1 - y)log(1 - p)
[0133] Among them, L Cls is the classification loss, L Res is the regression loss, λ j is the weight of the j-th sub-regression loss, n is the number of regressors, L Focal is the focal loss, α is the weight, p is the classification result, and y is the labeled classification label.
[0134] The comprehensive loss in step S240 is calculated from the classification loss and the segmentation loss. If the regression loss is not used, the comprehensive loss is:
[0135] L total =L focal +L Seg
[0136] Among them, L total is the comprehensive loss, L focal is the classification loss without the regression loss, L Seg is the segmentation loss.
[0137] If the regression loss is added, the comprehensive loss is:
[0138] L total =L Cls +L Seg
[0139] Among them, L total is the comprehensive loss, L Cls is the classification loss including the regression loss, L Seg is the segmentation loss.
[0140] In a possible implementation, refer to Figure 14, the multi-task model includes a shared encoder, a segmentation network, and a classification network. The classification network includes at least seven regressors, a fully connected layer, and a classifier. The shared encoder includes five encoding blocks. The segmentation network includes a multi-scale feature fusion self-attention module and a decoder. The decoder includes five decoding blocks. In one training, the cropped image samples are input into the shared encoder to obtain four feature maps of different scales, which are the outputs of the first four encoding blocks and the high-order feature map of the highest-order semantics output by the fifth encoding block. The four feature maps of different scales are fused through the multi-scale feature fusion self-attention module, and the fused features are restored to the corresponding positions and input into the corresponding decoder. The high-order feature map output by the fifth encoding block is directly input into the fifth decoding block. After being processed by the decoder, the segmentation result of the cropped image sample is obtained. The morphological features of the segmentation result are extracted, including the lesion area, the area of the smallest convex polygon corresponding to the lesion, the major axis and minor axis of the ellipse fitted to the lesion area, the number of contour pixels (perimeter) of the lesion, the angle (orientation) between the major axis of the fitted ellipse and the X-axis, and the diameter of the circle equal to the lesion area (equivalent diameter). The high-order feature map output by the fifth encoding block is simultaneously input into the seven regressors of the classification network, and is encoded into seven different 1×256-dimensional first vectors respectively. In one branch, each first vector is respectively input into the fully connected layer to obtain the predicted values of the respective first vectors. Based on the regression loss function, the regression loss is calculated according to the predicted values and the respective morphological sub-features, and the regression of different morphological feature values is performed. In another branch, the seven first vectors are concatenated into a 1×1792-dimensional second vector in this dimension, and then the second vector is fed into the DNN network for benign and malignant classification of gallbladder polyps.
[0141] During the training process, a dynamic pseudo-ground truth strategy is adopted to improve the regression effect of morphological feature values, that is, the seven major features of gallbladder polyps are extracted from the segmentation results segmented by the segmentation model in each iteration round as the pseudo-ground truth for the regression of feature values. Since the segmentation model is continuously optimized, the pseudo-ground truth of the feature values extracted by it is also continuously optimized. In this process, the regression effect of the feature values can also inversely optimize the shared encoder of the classification model and the segmentation model. Therefore, the morphological feature prior will also act on the segmentation process to further improve the segmentation effect. The pseudo-ground truth becomes more and more accurate with the continuous optimization of the segmentation results, promoting the mutual positive cycle improvement of the classification task and the segmentation task.
[0142] Furthermore, the embodiment of the present invention also provides an image processing method. Refer to Figure 15 , the method includes the following steps:
[0143] S310: Obtain the cropped image to be processed.
[0144] S320: Input the image to be processed for cropping into the multi-task model to obtain the segmentation result and classification result of the image to be processed for cropping. The multi-task model is trained by the above-mentioned multi-task model training method.
[0145] By using the multi-task model trained by the above-mentioned multi-task model training method to process the image to be processed for cropping, a more accurate segmentation result and classification result can be obtained.
[0146] In summary, for the multi-task model training method, image processing method, electronic device and storage medium provided by the embodiments of the present invention, through the shared encoder of the segmentation network and the classification network, their respective losses are calculated, and training is carried out according to the mutual guidance of the segmentation loss and the classification loss, so as to improve the classification accuracy of the model. The morphological features extracted from the output result of the segmentation model are used as the pseudo ground truth to optimize the regression model. The pseudo ground truth becomes more accurate as the segmentation result is continuously optimized, promoting a virtuous cycle of improvement between the classification task and the segmentation task. Through the multi-scale feature fusion self-attention mechanism, the multi-scale gallbladder polyp ultrasound image feature maps are extracted from the encoder for fusion, and the self-attention mechanism is used to model the long-range characteristics of the fusion features, enhancing the mutual utilization between different scale features.
[0147] In the embodiments provided by the present invention, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the device, method and computer program product according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0148] In addition, in each embodiment of the present invention, the functional modules can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.
[0149] If a function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a computer-readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0150] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
[0151] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to embrace all changes within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for training a multi-task model, characterized in that, The multi-task model includes a shared encoder, a segmentation network, and a classification network. The method includes: Inputting the cropped image sample into the shared encoder to obtain a first feature map. The cropped image sample includes an annotated classification label and an annotated segmentation label. Inputting the first feature map into the segmentation network to obtain a segmentation result, and calculating a segmentation loss according to the segmentation result and the annotated segmentation label. Inputting the first feature map into the classification network to obtain a classification result, and calculating a classification loss according to the classification result and the annotated classification label. Calculating a comprehensive loss of the multi-task model according to the classification loss and the segmentation loss, and updating the parameters of the multi-task model using the comprehensive loss.
2. The method according to claim 1, wherein The classification network includes at least one regressor and a classifier. The step of inputting the first feature map into the classification network to obtain a classification result and calculating a classification loss according to the classification result and the annotated classification label includes: Inputting the first feature map into each of the regressors to be encoded into first vectors respectively. Concatenating the first vectors in dimension to obtain a second vector, and inputting the second vector into the classifier to obtain the classification result of the cropped image sample. Calculating a classification loss based on a classification loss function according to the classification result and the annotated classification label.
3. The method according to claim 1, characterized in that The classification network includes at least one regressor, a fully connected layer, and a classifier. The step of inputting the first feature map into the classification network to obtain a classification result and calculating a classification loss according to the classification result and the annotated classification label includes: Extracting the morphological features of the segmentation result. The morphological features include morphological sub-features with the same number as the number of regressors. Inputting the first feature map into each of the regressors to be encoded into first vectors respectively. Inputting each of the first vectors into the fully connected layer to obtain predicted values of each of the first vectors respectively. Calculating a regression loss based on a regression loss function according to the predicted values and the morphological sub-features. Concatenating the first vectors in dimension to obtain a second vector, and inputting the second vector into the classifier to obtain the classification result of the cropped image sample. Calculating a classification loss based on a classification loss function according to the classification result, the annotated classification label, and the regression loss.
4. The method according to claim 3, wherein The regression loss function is: Among them, y i is the morphological sub-feature, y is the predicted value, δ is the hyperparameter, L Huber is the sub-regression loss of a single regressor, L Res is the regression loss, λ j is the weight of the j-th sub-regression loss, and n is the number of the regressors.
5. The method according to claim 1, characterized in that, The shared encoder includes at least three encoding blocks. The first feature map includes at least two second feature maps and a high-order feature map. The high-order feature map is the feature map output by the last encoding block. The segmentation network includes a multi-scale feature fusion self-attention module and a decoder. The decoder includes decoding blocks with the same number as each of the encoding blocks. The step of inputting the first feature map into the segmentation network to obtain a segmentation result and calculating a segmentation loss according to the segmentation result and the annotated segmentation label includes: Inputting each of the second feature maps into the multi-scale feature fusion self-attention module to obtain each third feature map encoded by self-attention. Input each of the third feature maps and the high-order feature map into the corresponding decoding block of the decoder to obtain the segmentation result of the cropped image sample; Based on the segmentation loss function, calculate the segmentation loss according to the segmentation result and the labeled segmentation label.
6. The method according to claim 5, wherein The step of inputting each of the second feature maps into the multi-scale feature fusion self-attention module to obtain each third feature map encoded by self-attention includes: Fuse each of the second feature maps to obtain a fused feature map; Divide the fused feature map into multiple window sub-features; Based on the position embedding formula, perform position embedding on each of the window sub-features to obtain the position-embedded features after position embedding; Perform self-attention encoding on the position-embedded features to obtain a self-attention feature map; Restore the position of the self-attention feature map to each third feature map corresponding to each of the second feature maps.
7. The method according to claim 6, wherein The position embedding formula is: Among them, z0 is the positional embedding feature, is the k-th window sub-feature, E is the hidden vector, E pos is the positional encoding matrix.
8. An image processing method, characterized in that, The method includes: Obtain a cropped image to be processed; Input the cropped image to be processed into a multi-task model to obtain the segmentation result and classification result of the cropped image to be processed; the multi-task model is trained by the multi-task model training method according to any one of claims 1 to 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
10. A storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by the processor, the method according to any one of claims 1 to 8 is implemented.