Image segmentation model training method, image segmentation method, device and electronic equipment
By introducing a discard module and an attention enhancement module in the image segmentation model, combining labeled and unlabeled images for iterative optimization, the problem of insufficient adaptive learning ability in the prior art is solved, and the performance and accuracy of the thyroid nodule segmentation task is significantly improved.
Patent Information
- Application Number
- CN202510200665.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The existing semi-supervised learning methods lack adaptive learning ability in the processing of thyroid nodules data, and it is difficult to accurately capture the morphology, texture and structural characteristics of the nodules, affecting the segmentation accuracy of the model in complex clinical scenarios.
By combining the discard module and the attention enhancement module, iterative optimization is performed using labeled and unlabeled images to automatically learn rich and essential features, and improve the adaptability and segmentation accuracy of the model under different data distributions.
The performance of the model in the thyroid nodule segmentation task is significantly improved, and the goals of nodules and artifacts, uneven areas can be better distinguished, and the robustness and training efficiency of the model are improved.
Smart Images

Figure CN119672043B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image segmentation, and in particular to an image segmentation model training method, an image segmentation method, a device and an electronic device. Background Art
[0002] Thyroid nodules are the most common clinical manifestation of thyroid diseases. Their accurate diagnosis and effective evaluation are crucial for the formulation of treatment strategies and the recovery of patients. Ultrasound examination is an important means of detecting thyroid nodules. Ultrasound images are obtained through ultrasound examination and segmented using nodule segmentation models.
[0003] However, training the model requires a large amount of labeled data, and the extreme scarcity of high-quality pixel-level labeled data has become a key bottleneck. In related technologies, semi-supervised learning methods can be used to train nodule segmentation models.
[0004] However, the consistency regularization method used in the existing semi-supervised learning method relies too much on artificially designed fixed strategies in the data enhancement stage, lacks the ability to adaptively learn the distribution characteristics of thyroid nodule data, and is difficult to accurately capture the essential characteristics of thyroid nodules in terms of morphology, texture, structure, etc., which in turn affects the segmentation accuracy of the model in complex clinical scenarios. Summary of the invention
[0005] In view of this, an object of an embodiment of the present invention is to provide an image segmentation model training method, an image segmentation method, an apparatus and an electronic device to at least partially improve the above-mentioned problem.
[0006] In order to achieve the above purpose, the technical solution adopted by the embodiment of the present invention is as follows:
[0007] In a first aspect, an embodiment of the present invention provides an image segmentation model training method, the method comprising:
[0008] Acquire a labeled image, an unlabeled image, an unlabeled weakly perturbed image, and an unlabeled strongly perturbed image; wherein the labeled image includes a real label, the unlabeled weakly perturbed image is obtained by weakly perturbing the unlabeled image, and the unlabeled strongly perturbed image is obtained by strongly perturbing the unlabeled image;
[0009] Inputting the labeled image and the unlabeled image into an image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively;
[0010] Inputting the unlabeled strongly disturbed image into the image segmentation model, and processing it through a discarding module to obtain a strongly disturbed prediction result; the discarding module is used to discard some features of the unlabeled strongly disturbed image;
[0011] Inputting the unlabeled weakly perturbed image into the image segmentation model, and processing it through an attention enhancement module to obtain a weakly perturbed prediction result; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image;
[0012] Loss information is calculated according to the strong disturbance prediction result, the weak disturbance prediction result, the unlabeled prediction result and the labeled prediction result, and the image segmentation model is iteratively optimized according to the loss information to obtain a mature image segmentation model.
[0013] Optionally, the image segmentation model includes an encoder, a decoder and an attention enhancement module, the attention enhancement module includes a residual convolution submodule and an attention submodule, and the unlabeled weakly perturbed image is input into the image segmentation model and processed by the attention enhancement module to obtain a weakly perturbation prediction result, including:
[0014] Inputting the unlabeled weakly perturbed image into the encoder for feature extraction to obtain a first feature map, and dividing the first feature map into a plurality of four-dimensional tensors of a preset batch size;
[0015] For each of the four-dimensional tensors, input the four-dimensional tensor into the residual convolution submodule, process it through a 1x1 convolution layer, generate a residual connection, and obtain a residual tensor;
[0016] Inputting the four-dimensional tensor into the attention submodule to obtain an attention feature map;
[0017] The attention feature map is added to the residual tensor and input into the decoder to obtain a weak perturbation prediction result.
[0018] Optionally, the attention submodule includes a channel attention layer, a spatial attention layer, and a dynamic convolution layer, and the inputting the four-dimensional tensor into the attention submodule to obtain an attention feature map includes:
[0019] Inputting the four-dimensional tensor into the channel attention layer to obtain a channel attention weight;
[0020] Multiply the channel attention weight by the four-dimensional tensor element by element to obtain a channel attention feature map;
[0021] Inputting the channel attention feature map into the spatial attention layer to obtain a spatial attention weight;
[0022] Multiplying the spatial attention weight by the channel attention feature map element by element to obtain a spatial attention feature map;
[0023] The spatial attention feature map is input into the dynamic convolution layer to obtain the final attention feature map.
[0024] Optionally, inputting the four-dimensional tensor into the channel attention layer to obtain a channel attention weight includes:
[0025] Performing average pooling and maximum pooling on the feature map in the four-dimensional tensor to obtain an average pooling feature map and a maximum pooling feature map;
[0026] Processing the average pooling feature map and the maximum pooling feature map through a fully connected layer to obtain an average pooling intermediate representation and a maximum pooling intermediate representation;
[0027] The average pooled intermediate representation and the maximum pooled intermediate representation are added and activated by an activation function to obtain a channel attention weight.
[0028] Optionally, inputting the channel attention feature map into the spatial attention layer to obtain the spatial attention weight includes:
[0029] Calculate the average value and maximum value of the channel attention feature map in the channel dimension to obtain an average feature map and a maximum feature map;
[0030] The average feature map and the maximum feature map are concatenated and processed by convolution and activation function to obtain the spatial attention weight.
[0031] Optionally, inputting the spatial attention feature map into the dynamic convolution layer to obtain a final attention feature map includes:
[0032] Predicting the dynamic convolution kernel of the dynamic convolution layer through a kernel prediction network;
[0033] The spatial attention feature map is reshaped using the dynamic convolution kernel to obtain a final attention feature map.
[0034] Optionally, the image segmentation model includes an encoder and a decoder, and the step of inputting the unlabeled strongly disturbed image into the image segmentation model and processing it through a discarding module to obtain a strongly disturbed prediction result includes:
[0035] Inputting the unlabeled strongly disturbed image into the encoder for feature extraction to obtain a second feature map;
[0036] Inputting the second feature map into the discarding module, randomly discarding some features, and obtaining a second feature map discarding some features;
[0037] The second feature map discarding some features is input into the decoder for upsampling to obtain a strong disturbance prediction result.
[0038] In a second aspect, an embodiment of the present invention provides an image segmentation method, the method comprising:
[0039] Obtain the image to be segmented;
[0040] The image to be segmented is input into an image segmentation model to obtain a segmented image; the image segmentation model is trained by the method described in any one of the first aspects.
[0041] In a third aspect, an embodiment of the present invention provides an image segmentation model training device, the device comprising:
[0042] An image acquisition unit, used to acquire an annotated image, an unannotated image, an unannotated weakly perturbed image, and an unannotated strongly perturbed image; wherein the annotated image includes a real annotation, the unannotated weakly perturbed image is obtained by weakly perturbing the unannotated image, and the unannotated strongly perturbed image is obtained by strongly perturbing the unannotated image;
[0043] A first prediction unit, used for inputting the labeled image and the unlabeled image into an image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively;
[0044] A second prediction unit is used to input the unlabeled strongly disturbed image into the image segmentation model, and obtain a strongly disturbed prediction result after being processed by a discarding module; the discarding module is used to discard some features of the unlabeled strongly disturbed image;
[0045] A third prediction unit is used to input the unlabeled weakly perturbed image into the image segmentation model, and obtain a weakly perturbed prediction result after being processed by an attention enhancement module; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image;
[0046] An iterative optimization unit is used to calculate loss information according to the strong disturbance prediction result, the weak disturbance prediction result, the unlabeled prediction result and the labeled prediction result, and iteratively optimize the image segmentation model according to the loss information to obtain a mature image segmentation model.
[0047] In a fourth aspect, an embodiment of the present invention provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements any of the above-described methods when executing the program.
[0048] An image segmentation model training method, image segmentation method, device and electronic device provided by the embodiments of the present invention can automatically learn rich features and essential features under different data distributions by combining a discarding module and an attention enhancement module. This feature ensures that the model can quickly adapt and optimize its segmentation effect when processing complex medical images. The attention enhancement module significantly improves the performance of the model in the thyroid nodule segmentation task by enhancing important features, improving the richness of feature representation, reducing redundant features, enhancing the robustness of the model and improving training efficiency. By combining intrinsic features with diversified features, the performance of the model in the thyroid nodule segmentation task is significantly improved. By utilizing the features of labeled data and the enhanced features of unlabeled data, a multi-level feature representation is formed, which enables the model to better distinguish nodules from artifacts, uneven areas and other targets.
[0049] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are specifically cited below and described in detail with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0051] Figure 1 A schematic structural block diagram of an electronic device provided by an embodiment of the present invention;
[0052] Figure 2 A flowchart of an image segmentation model training method provided by an embodiment of the present invention;
[0053] Figure 3 A schematic structural block diagram of an image segmentation model provided by an embodiment of the present invention;
[0054] Figure 4 A data flow diagram of an unlabeled weakly disturbed image provided by an embodiment of the present invention;
[0055] Figure 5 A schematic diagram of a flow chart of step S240 provided in an embodiment of the present invention;
[0056] Figure 6 A schematic structural block diagram of an attention enhancement module provided by an embodiment of the present invention;
[0057] Figure 7 A schematic diagram of a flow chart of step S243 provided in an embodiment of the present invention;
[0058] Figure 8 A schematic diagram of a flow chart of an image segmentation method provided by an embodiment of the present invention;
[0059] Fig. 9 A schematic structural block diagram of an image segmentation model training device provided in an embodiment of the present invention.
[0060] Icon: 100-electronic device; 101-memory; 102-communication interface; 103-processor; 104-communication bus; 30-image segmentation model; 31-encoder; 32-decoder; 33-attention enhancement module; 331-residual convolution submodule; 332-attention submodule; 3321-channel attention layer; 3322-spatial attention layer; 3323-dynamic convolution layer; 34-discarding module; 400-image segmentation model training device; 410-image acquisition unit; 420-first prediction unit; 430-second prediction unit; 440-third prediction unit; 450-iterative optimization unit. DETAILED DESCRIPTION
[0061] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0062] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0063] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0064] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0065] With the continuous development of medical science, the incidence of thyroid diseases has gradually increased and has become one of the focuses of global public health attention. As the most common clinical manifestation of thyroid disease, the accurate diagnosis and effective evaluation of thyroid nodules are crucial for the formulation of treatment strategies and the recovery of patients. Ultrasound examination occupies an important position in the initial screening and daily monitoring of thyroid diseases due to its real-time, convenient and radiation-free advantages.
[0066] However, the problems of blurred boundaries, tissue overlap, and artifact interference in ultrasound images make automatic segmentation algorithms based on computer vision technology face huge challenges in practical applications. These technical barriers limit the accuracy and stability of the algorithm, especially in the training stage of deep learning models, where the extreme scarcity of high-quality pixel-level annotated data becomes a key bottleneck. In addition, the annotation of medical images requires personnel with deep professional knowledge, which is time-consuming and labor-intensive, further exacerbating the shortage of high-quality data.
[0067] Using semi-supervised learning methods to train the model can solve some of the problems, but the consistency regularization method used in the existing semi-supervised learning method relies too much on artificially designed fixed strategies in the data enhancement stage, lacks the ability to adaptively learn the distribution characteristics of thyroid nodule data, and is difficult to accurately capture the essential characteristics of thyroid nodules in terms of morphology, texture, structure, etc., which in turn affects the segmentation accuracy of the model in complex clinical scenarios.
[0068] Based on the above situation, the embodiment of the present invention provides an image segmentation model training method, an image segmentation method, a device and an electronic device, by inputting annotated images, unlabeled images, unlabeled strongly disturbed images and unlabeled weakly disturbed images into the image segmentation model, the unlabeled strongly disturbed images are processed by the discard module, the unlabeled weakly disturbed images are processed by the attention enhancement module, the loss information is calculated according to each output result, and the image segmentation model is iteratively optimized according to the loss information to obtain a mature image segmentation model. By combining the discard module and the attention enhancement module, rich features and important features can be automatically learned under different data distributions, thereby improving the segmentation accuracy of the model in complex clinical scenarios.
[0069] To implement the process steps and functions of each example of the present invention, please refer to Figure 1 , Figure 1 A schematic structural block diagram of an electronic device provided in an embodiment of the present invention. The electronic device 100 includes a memory 101 and a processor 103, and the memory 101 and the processor 103 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses 104 or signal lines. The memory 101 can be used to store software programs and modules, and the processor 103 executes the software programs and modules stored in the memory 101, thereby performing various functional applications and data processing.
[0070] The electronic device 100 may be, but is not limited to, a personal computer (PC), a server, a distributed computer, etc. It is understandable that the electronic device 100 is not limited to a physical server, but may also be a virtual machine on a physical server, a virtual machine built on a cloud platform, etc., which can provide the same functions as the server or virtual machine. The operating system of the electronic device 100 may be, but is not limited to, a Windows system, a Linux system, etc.
[0071] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc.
[0072] The communication connection between the electronic device 100 and an external device is achieved through at least one communication interface 102 (which can be wired or wireless).
[0073] The processor 103 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the embodiment of the present invention can be completed by the hardware integrated logic circuit in the processor 103 or the instructions in the form of software. The processor 103 can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components.
[0074] Understandably, Figure 1 The structure shown is for illustration only. The electronic device 100 may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0075] The following is an exemplary description of the image segmentation model training method provided by the present invention. Specifically, Figure 2 A flowchart of an image segmentation model training method provided by an embodiment of the present invention is shown in FIG. Figure 2 , the execution subject of this method can be the above Figure 1 The electronic device 100 shown in FIG. 1 includes: Figure 2 The following steps are shown:
[0076] S210: Obtain an annotated image, an unannotated image, an unannotated weakly perturbed image, and an unannotated strongly perturbed image; wherein the annotated image includes a real annotation, the unannotated weakly perturbed image is obtained by weakly perturbing the unannotated image, and the unannotated strongly perturbed image is obtained by strongly perturbing the unannotated image.
[0077] Annotated images are images with real annotations. Annotators with profound professional medical knowledge annotate the boundaries, types, and surrounding tissues of thyroid nodules in thyroid ultrasound images.
[0078] The unlabeled images can directly use the thyroid ultrasound images. Preferably, the thyroid ultrasound images can be further screened to remove noise.
[0079] Unlabeled weakly perturbed images apply mild random transformations such as rotation, scaling, translation, etc. to unlabeled images to create new training samples, but these transformations will not significantly change the essential characteristics of the image.
[0080] Unlabeled strongly perturbed images apply large random transformations to unlabeled images, such as random cropping, adding Gaussian noise, changing brightness / contrast, etc., so that the image undergoes large changes but still retains certain key features.
[0081] S220: Input the labeled image and the unlabeled image into the image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively.
[0082] Among them, the backbone segmentation network of the image segmentation model can use the UNet architecture, that is, the structure of the encoder and decoder.
[0083] See also Figure 3 The unlabeled image without data enhancement is input into the image segmentation model 30, and directly passes through the encoder 31 and the decoder 32 to obtain the unlabeled prediction result, which is subsequently used as the pseudo label of the unlabeled weakly perturbed image and the unlabeled strongly perturbed image.
[0084] The labeled images are input into the image segmentation model to obtain labeled prediction results, which will be directly compared with the true annotations to guide the training of the model.
[0085] S230: Inputting the unlabeled strongly disturbed image into the image segmentation model, and processing it through a discarding module to obtain a strongly disturbed prediction result; the discarding module processing is used to discard some features of the unlabeled strongly disturbed image.
[0086] For unlabeled strongly disturbed images, after being input into the image segmentation model to extract features, they are processed by the discard module 34. According to the discard probability, the outputs of some neurons will be randomly set to zero. The image segmentation model uses the remaining valid features for prediction to obtain the strongly disturbed prediction results.
[0087] S240: Input the unlabeled weakly perturbed image into the image segmentation model, and process it through the attention enhancement module 33 to obtain a weakly perturbed prediction result; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image.
[0088] For unlabeled weakly perturbed images, after the features are extracted by the input image segmentation model, they are processed by the attention enhancement module to obtain the weakly perturbated prediction results. The attention enhancement module can extract some important information from the features and retain the key intrinsic features while removing redundant features.
[0089] S250: Calculate loss information based on the strong perturbation prediction results, the weak perturbation prediction results, the unlabeled prediction results, and the labeled prediction results, iteratively optimize the image segmentation model based on the loss information, and obtain a mature image segmentation model. The loss information includes first loss information calculated from the labeled prediction results and the labeled image, second loss information calculated from the weak perturbation prediction results and the unlabeled prediction results, and third loss information calculated from the strong perturbation prediction results and the unlabeled prediction results.
[0090] Based on the four prediction results obtained, the loss information is calculated, where the loss function aims to effectively utilize the diverse and intrinsic characteristics of labeled and unlabeled data.
[0091] See also Figure 3 , the input of the loss function includes the labeled image , unlabeled images , unlabeled weakly perturbed images And unlabeled strongly perturbed images .
[0092] For annotated images , its loss is defined as the standard cross entropy loss, and the first loss information calculated by the labeled prediction result and the labeled image is expressed as:
[0093]
[0094] in, is the sample size, It is a true mark. There are labeled prediction results, is the Dice loss function.
[0095] For unlabeled images , the pseudo labels generated by the image segmentation model To conduct training.
[0096] For unlabeled weakly perturbed images , its loss function also adopts the cross entropy form, and the second loss information calculated by the weak perturbation prediction result and the unlabeled prediction result is expressed as:
[0097]
[0098] in, is the number of unlabeled weakly perturbed image samples, It is the weak perturbation prediction result.
[0099] For unlabeled strongly perturbed images , the third loss information calculated by the strong perturbation prediction result and the unlabeled prediction result is expressed as:
[0100]
[0101] in, It is the strong disturbance prediction result.
[0102] Combining all the above losses, the total loss function is defined as:
[0103]
[0104] in, , and is a hyperparameter used to balance the impact of various loss terms.
[0105] The loss information is calculated by the loss function, and the image segmentation model is iteratively optimized according to the loss information. After reaching the preset training conditions, a mature image segmentation model is obtained. The image segmentation model is used to segment thyroid ultrasound images, which can improve the segmentation accuracy of thyroid nodules.
[0106] By combining the discard module and the attention enhancement module, this method can automatically learn rich features and important features under different data distributions, thereby improving the segmentation accuracy of the model in complex clinical scenarios.
[0107] In one possible implementation, in order to extract important information from unlabeled weakly perturbed images, see Figure 4 The image segmentation model may include an encoder 31, a decoder 32, and an attention enhancement module 33. The attention enhancement module 33 includes a residual convolution submodule 331 and an attention submodule 332. Figure 5 , the above step S240 may include the following steps:
[0108] S241: Input the unlabeled weakly perturbation image into the encoder for feature extraction to obtain a first feature map, and divide the first feature map into a plurality of four-dimensional tensors of a preset batch size.
[0109] It can be understood that the unlabeled weakly perturbed image is a batch of images. The first feature map obtained by extracting the unlabeled weakly perturbed image through the encoder is divided into batches, and each batch is a four-dimensional tensor, which is expressed as ,in is the batch size, that is, the preset batch size, is the number of channels, and are the height and width of the first feature map respectively.
[0110] S242: For each four-dimensional tensor, the four-dimensional tensor is input into the residual convolution submodule, processed through a 1x1 convolution layer, a residual connection is generated, and a residual tensor is obtained.
[0111] Each four-dimensional tensor is processed separately and input into the residual convolution submodule through a The convolutional layer is processed to generate residual connections to form a residual tensor. The process can be expressed as:
[0112]
[0113] in, is a four-dimensional tensor, is the residual tensor.
[0114] S243: Input the four-dimensional tensor into the attention submodule to obtain the attention feature map.
[0115] S244: Add the attention feature map to the residual tensor and input it into the decoder to obtain the weak perturbation prediction result.
[0116] At the same time, the four-dimensional tensor is input into another module, the attention submodule, which is used to identify which areas in the image are more important or should be paid more attention to. After processing, an attention feature map is obtained. Next, the attention feature map generated by the attention submodule is added to the residual tensor generated by the residual convolution submodule, and the result of the addition is input into the decoder to finally obtain the prediction result of the corresponding unlabeled weak perturbation image.
[0117] In one possible implementation, in order to extract richer intrinsic features from unlabeled weakly perturbed images, see Figure 6 The attention submodule may include a channel attention layer 3321, a spatial attention layer 3322, and a dynamic convolution layer 3323, see Figure 7 , the above step S243 may include the following steps:
[0118] S2431: Input the four-dimensional tensor into the channel attention layer to obtain the channel attention weight.
[0119] There are many ways to calculate the channel attention weight. In one possible implementation, the step of calculating the channel attention weight may include:
[0120] S24311: Perform average pooling and maximum pooling on the feature map in the four-dimensional tensor to obtain an average pooling feature map and a maximum pooling feature map.
[0121] S24312: The average pooling feature map and the maximum pooling feature map are processed through a fully connected layer to obtain an average pooling intermediate representation and a maximum pooling intermediate representation.
[0122] S24313: Add the average pooled intermediate representation and the maximum pooled intermediate representation, and activate them through the activation function to obtain the channel attention weight.
[0123] The calculation steps of the channel attention weight can be expressed by the following formula:
[0124]
[0125] in, is a four-dimensional tensor, represents average pooling, represents the maximum pooling, represents the fully connected layer processing, Represents the activation function, which can be a sigmoid activation function. is the channel attention weight.
[0126] S2432: Multiply the channel attention weight by the four-dimensional tensor element by element to obtain the channel attention feature map.
[0127] By multiplying the channel attention weights by the four-dimensional tensor element by element, the features of important channels can be enhanced. The calculation formula is as follows:
[0128]
[0129] S2433: Input the channel attention feature map into the spatial attention layer to obtain the spatial attention weight.
[0130] There are many ways to calculate the spatial attention weight. In one possible implementation, the steps of calculating the spatial attention weight may include:
[0131] S24331: Calculate the average and maximum values of the channel attention feature map in the channel dimension to obtain the average feature map and the maximum feature map.
[0132] S24332: Concatenate the average feature map and the maximum feature map, and process them through convolution and activation functions to obtain the spatial attention weight.
[0133] The calculation steps of the spatial attention weight can be expressed by the following formula:
[0134]
[0135] in, is the channel attention feature map, represents the average value, Indicates the maximum value, Indicates splicing, represents the convolution process, Represents the activation function, which can be a sigmoid activation function. is the spatial attention weight.
[0136] S2434: Multiply the spatial attention weight by the channel attention feature map element by element to obtain the spatial attention feature map.
[0137] By multiplying the spatial attention weights by the channel attention feature map element by element, the features of the spatially important regions can be further enhanced. The calculation formula is as follows:
[0138]
[0139] S2435: Input the spatial attention feature map into the dynamic convolution layer to obtain the final attention feature map.
[0140] First, the dynamic convolution layer predicts the dynamic convolution kernel of the dynamic convolution layer through the kernel prediction network. The prediction formula is as follows:
[0141]
[0142] in, is the kernel prediction network, represents the dimension of the predicted dynamic convolution kernel K, B is the batch size, i.e. the number of data samples processed at a time, is the number of output channels, that is, the number of channels of the spatial attention feature map, K is the size of the convolution kernel, Indicates the number of weights of the convolution kernel corresponding to each output channel, Indicates that the height and width of the convolution kernel are both 1.
[0143] The spatial attention feature map is reshaped using a dynamic convolution kernel to obtain the final attention feature map.
[0144] The dynamic convolution operation is as follows:
[0145]
[0146] in, is the dynamic convolution, is the spatial attention feature map, is the dynamic convolution kernel, is the final output attention feature map.
[0147] In a possible implementation, the above step S230 may include the following steps:
[0148] S231: Input the unlabeled strongly perturbed image into the encoder for feature extraction to obtain a second feature map.
[0149] S232: Input the second feature map into a discarding module, randomly discard some features, and obtain a second feature map with some features discarded.
[0150] S233: Input the second feature map with some features discarded into the decoder for upsampling to obtain a strong disturbance prediction result.
[0151] The unlabeled strongly perturbed image is input into the encoder, which will automatically extract the key features in the image and generate a low-dimensional feature representation, called the second feature map. Some features are randomly discarded through a discarding module (such as the Dropout layer). This can prevent the model from overfitting and enhance its generalization ability. The discarded second feature map is then input into the decoder. The decoder is responsible for upsampling the low-resolution feature map back to a high-resolution image or feature representation, and finally obtaining the prediction result of the strongly perturbed image.
[0152] Furthermore, based on the above-mentioned image segmentation model training method, the embodiment of the present invention also provides an image segmentation method, referring to Figure 8 , the method comprises the following steps:
[0153] S310: Obtain an image to be segmented.
[0154] The image to be segmented may be a thyroid ultrasound image obtained by ultrasound scanning of the patient.
[0155] S320: Inputting the image to be segmented into an image segmentation model to obtain a segmented image; wherein the image segmentation model is trained by the above-mentioned image segmentation model training method.
[0156] By inputting the image to be segmented into a trained image segmentation model, nodules and the like in the image to be segmented can be accurately segmented, and targets such as nodules and artifacts in the image can be distinguished.
[0157] Furthermore, the present invention also provides an image segmentation model training device, see Fig. 9 , the image segmentation model training device 400 includes:
[0158] The image acquisition unit 410 is used to acquire annotated images, unlabeled images, unlabeled weakly perturbed images, and unlabeled strongly perturbed images; wherein the annotated images include real annotations, the unlabeled weakly perturbed images are obtained by weakly perturbing the unlabeled images, and the unlabeled strongly perturbed images are obtained by strongly perturbing the unlabeled images.
[0159] The first prediction unit 420 is used to input the labeled image and the unlabeled image into the image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively.
[0160] The second prediction unit 430 is used to input the unlabeled strongly disturbed image into the image segmentation model, and obtain a strongly disturbed prediction result after being processed by the discarding module; the discarding module processing is used to discard some features of the unlabeled strongly disturbed image.
[0161] The third prediction unit 440 is used to input the unlabeled weakly perturbed image into the image segmentation model, and obtain the weakly perturbed prediction result after being processed by the attention enhancement module; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image.
[0162] The iterative optimization unit 450 is used to calculate the loss information according to the strong perturbation prediction result, the weak perturbation prediction result, the unlabeled prediction result and the labeled prediction result, and iteratively optimize the image segmentation model according to the loss information to obtain a mature image segmentation model. The loss information includes the first loss information calculated by the labeled prediction result and the labeled image, the second loss information calculated by the weak perturbation prediction result and the unlabeled prediction result, and the third loss information calculated by the strong perturbation prediction result and the unlabeled prediction result.
[0163] In summary, an image segmentation model training method, an image segmentation method, a device and an electronic device provided by an embodiment of the present invention can automatically learn rich features and essential features under different data distributions by combining a discarding module and an attention enhancement module. This feature ensures that the model can quickly adapt and optimize its segmentation effect when processing complex medical images. The attention enhancement module significantly improves the performance of the model in the thyroid nodule segmentation task by enhancing important features, improving the richness of feature representation, reducing redundant features, enhancing the robustness of the model and improving training efficiency. By combining intrinsic features with diversified features, the performance of the model in the thyroid nodule segmentation task is significantly improved. By utilizing the features of labeled data and the enhanced features of unlabeled data, a multi-level feature representation is formed, which enables the model to better distinguish between nodules and artifacts, uneven areas and other targets.
[0164] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a module, a program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0165] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0166] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a computer-readable storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0167] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
[0168] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A method for training an image segmentation model, characterized in that: The method comprises: Acquire a labeled image, an unlabeled image, an unlabeled weakly perturbed image, and an unlabeled strongly perturbed image; wherein the labeled image includes a real label, the unlabeled weakly perturbed image is obtained by weakly perturbing the unlabeled image, and the unlabeled strongly perturbed image is obtained by strongly perturbing the unlabeled image; Inputting the labeled image and the unlabeled image into an image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively; Inputting the unlabeled strongly disturbed image into the image segmentation model, and processing it through a discarding module to obtain a strongly disturbed prediction result; the discarding module is used to discard some features of the unlabeled strongly disturbed image; Inputting the unlabeled weakly perturbed image into the image segmentation model, and processing it through an attention enhancement module to obtain a weakly perturbed prediction result; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image; The loss information is calculated according to the strong perturbation prediction result, the weak perturbation prediction result, the unlabeled prediction result and the labeled prediction result, and the image segmentation model is iteratively optimized according to the loss information to obtain a mature image segmentation model; wherein the loss information includes first loss information calculated by the labeled prediction result and the labeled image, second loss information calculated by the weak perturbation prediction result and the unlabeled prediction result, and third loss information calculated by the strong perturbation prediction result and the unlabeled prediction result.
2. The method according to claim 1, characterized in that The image segmentation model includes an encoder, a decoder and an attention enhancement module, the attention enhancement module includes a residual convolution submodule and an attention submodule, the unlabeled weak perturbation image is input into the image segmentation model, and is processed by the attention enhancement module to obtain a weak perturbation prediction result, including: Inputting the unlabeled weakly perturbed image into the encoder for feature extraction to obtain a first feature map, and dividing the first feature map into a plurality of four-dimensional tensors of a preset batch size; For each of the four-dimensional tensors, input the four-dimensional tensor into the residual convolution submodule, process it through a 1x1 convolution layer, generate a residual connection, and obtain a residual tensor; Inputting the four-dimensional tensor into the attention submodule to obtain an attention feature map; The attention feature map is added to the residual tensor and input into the decoder to obtain a weak perturbation prediction result.
3. The method according to claim 2, characterized in that The attention submodule includes a channel attention layer, a spatial attention layer and a dynamic convolution layer. The four-dimensional tensor is input into the attention submodule to obtain an attention feature map, including: Inputting the four-dimensional tensor into the channel attention layer to obtain a channel attention weight; Multiply the channel attention weight by the four-dimensional tensor element by element to obtain a channel attention feature map; Inputting the channel attention feature map into the spatial attention layer to obtain a spatial attention weight; Multiplying the spatial attention weight by the channel attention feature map element by element to obtain a spatial attention feature map; The spatial attention feature map is input into the dynamic convolution layer to obtain the final attention feature map.
4. The method according to claim 3, characterized in that The step of inputting the four-dimensional tensor into the channel attention layer to obtain the channel attention weight includes: Performing average pooling and maximum pooling on the feature map in the four-dimensional tensor to obtain an average pooling feature map and a maximum pooling feature map; Processing the average pooling feature map and the maximum pooling feature map through a fully connected layer to obtain an average pooling intermediate representation and a maximum pooling intermediate representation; The average pooled intermediate representation and the maximum pooled intermediate representation are added and activated by an activation function to obtain a channel attention weight.
5. The method according to claim 3, characterized in that: The step of inputting the channel attention feature map into the spatial attention layer to obtain the spatial attention weight comprises: Calculate the average value and maximum value of the channel attention feature map in the channel dimension to obtain an average feature map and a maximum feature map; The average feature map and the maximum feature map are concatenated and processed by convolution and activation function to obtain the spatial attention weight.
6. The method according to claim 3, characterized in that The step of inputting the spatial attention feature map into the dynamic convolution layer to obtain a final attention feature map comprises: Predicting the dynamic convolution kernel of the dynamic convolution layer through a kernel prediction network; The spatial attention feature map is reshaped using the dynamic convolution kernel to obtain a final attention feature map.
7. The method according to any one of claims 1 to 6, characterized in that: The image segmentation model includes an encoder and a decoder. The unlabeled strongly disturbed image is input into the image segmentation model and processed by a discarding module to obtain a strongly disturbed prediction result, including: Inputting the unlabeled strongly disturbed image into the encoder for feature extraction to obtain a second feature map; Inputting the second feature map into the discarding module, randomly discarding some features, and obtaining a second feature map discarding some features; The second feature map discarding some features is input into the decoder for upsampling to obtain a strong disturbance prediction result.
8. An image segmentation method, characterized in that: The method comprises: Obtain the image to be segmented; The image to be segmented is input into an image segmentation model to obtain a segmented image; the image segmentation model is trained by the method according to any one of claims 1 to 7.
9. An image segmentation model training device, characterized in that: The device comprises: An image acquisition unit, used to acquire an annotated image, an unannotated image, an unannotated weakly perturbed image, and an unannotated strongly perturbed image; wherein the annotated image includes a real annotation, the unannotated weakly perturbed image is obtained by weakly perturbing the unannotated image, and the unannotated strongly perturbed image is obtained by strongly perturbing the unannotated image; A first prediction unit, used for inputting the labeled image and the unlabeled image into an image segmentation model to obtain a labeled prediction result and an unlabeled prediction result respectively; A second prediction unit is used to input the unlabeled strongly disturbed image into the image segmentation model, and obtain a strongly disturbed prediction result after being processed by a discarding module; the discarding module is used to discard some features of the unlabeled strongly disturbed image; A third prediction unit is used to input the unlabeled weakly perturbed image into the image segmentation model, and obtain a weakly perturbed prediction result after being processed by an attention enhancement module; the attention enhancement module is used to extract the intrinsic features of the unlabeled weakly perturbed image; An iterative optimization unit is used to calculate loss information according to the strong perturbation prediction result, the weak perturbation prediction result, the unlabeled prediction result and the labeled prediction result, and iteratively optimize the image segmentation model according to the loss information to obtain a mature image segmentation model; wherein the loss information includes first loss information calculated by the labeled prediction result and the labeled image, second loss information calculated by the weak perturbation prediction result and the unlabeled prediction result, and third loss information calculated by the strong perturbation prediction result and the unlabeled prediction result.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Pathological image segmentation method and system fusing global and local similarity measurement
CN117523187A
Semi-supervised medical image segmentation method for mixing data disturbance and feature enhanced disturbance
CN118628739A