Method and apparatus for segmenting organ images

By introducing global and local feature fusion units into the U-Net network, the target organ segmentation prediction model is constructed, which solves the robustness and accuracy of OAR image outline in radiation therapy, and achieves efficient and accurate organ image segmentation.

CN120107292BActive Publication Date: 2025-07-22QINGDAO UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510591593.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-07-22
Estimated Expiration
2045-05-09

AI Technical Summary

Technical Problem

Prior Art In radiotherapy, the outline of radiotherapy-threatening organ (OAR) images is poorly robust and have low accuracy. Especially in the context of low contrast and complexity, traditional U-Net networks have problems such as insufficient feature extraction and local edge errors.

Method used

Introducing global and local feature fusion units in the U-Net network structure, including hollow convolutional layer, attention-weighted block and batch normalization layer, the target organ segmentation prediction model is built, and through multi-scale feature fusion and attention-weighted block, the key regional features are enhanced, irrelevant features are reduced, and the number of model parameters is reduced.

Benefits of technology

It realizes that the number of model parameters is reduced while ensuring high accuracy, improves the robustness and accuracy of organ image segmentation, solves the shortcomings of outlines in traditional methods, and is suitable for resource-constrained environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107292B_ABST
    Figure CN120107292B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for segmenting organ images, relating to the technical field of image processing. A method for segmenting organ images includes: constructing a target organ segmentation annotation dataset, where the target organ segmentation annotation dataset includes multiple groups of target organ image samples with segmentation annotations; constructing an initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture; training the initial segmentation model using the target organ segmentation annotation dataset to obtain a target organ segmentation prediction model; and inputting the image data of the target organ to be segmented into the target organ segmentation prediction model to obtain a target organ simulated segmentation result. According to the technical solution of the embodiments of the present application, a global and local feature fusion unit is introduced into the U-Net network structure to construct a target organ segmentation prediction model, which reduces the number of model parameters while ensuring high accuracy, and realizes the segmentation of organ images with good robustness and high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly relates to a method and device for segmenting organ images. Background Art

[0002] Currently, radiotherapy (RT) is one of the main treatment methods for various cancers. Accurate dose distribution is crucial for improving treatment effectiveness and protecting normal tissues. The delineation of radiotherapy organs at risk (OAR) images is the core technology for accurate dose distribution. However, there are various radiotherapy organs at risk. Some OARs have a high similarity in CT gray scale with surrounding soft tissues, blood vessels and other structures, and their morphological changes are diverse. Traditional manual delineation has poor robustness, low accuracy and low efficiency.

[0003] To reduce the above deficiencies, many researchers currently try to use deep learning (DL) algorithms to automatically segment OARs. U-Net is a network structure widely used in medical image segmentation. However, for the cases of low contrast, complex background and variable organ morphology, traditional U-Net still has problems such as insufficient feature extraction and local edge errors.

[0004] In summary, the related technologies have problems of poor robustness and low accuracy. Summary of the Invention

[0005] Based on this, this application provides a method and device for segmenting organ images, realizing the segmentation of organ images with good robustness and high accuracy.

[0006] According to one aspect of this application, a method for segmenting organ images is proposed, including: constructing a target organ segmentation annotation data set, where the target organ segmentation annotation data set includes multiple groups of target organ image samples with segmentation annotations; constructing an initial segmentation model based on a pre-constructed global and local feature fusion unit and U-Net network architecture; training the initial segmentation model using the target organ segmentation annotation data set to obtain a target organ segmentation prediction model; inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain a target organ simulated segmentation result.

[0007] According to some embodiments, the global and local feature fusion unit includes a dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a fully convolutional layer, a second attention weighting block and a second batch normalization layer connected in sequence, or a fully convolutional layer, a first attention weighting block, a first batch normalization layer, a dilated convolutional layer, a second attention weighting block and a second batch normalization layer connected in sequence.

[0008] According to some embodiments, an initial segmentation model is constructed based on a pre-constructed global and local feature fusion unit and a U-Net network architecture, including: splicing a plurality of global and local feature fusion units to construct a multi-scale global and local feature fusion module; using the multi-scale global and local feature fusion module as an encoding layer, and connecting a plurality of encoding layers based on the U-Net network architecture to form an encoder; based on the U-Net network architecture, connecting an input layer, the encoder, a pre-constructed decoder, and an output layer to obtain the initial segmentation model.

[0009] According to some embodiments, constructing the initial segmentation model based on the pre-constructed global and local feature fusion unit and the U-Net network architecture further includes: residually connecting the global and local feature fusion unit to the input layer.

[0010] According to some embodiments, the method further includes: the dilated convolutional layer and the fully convolutional layer of each global and local feature fusion unit run in parallel; and / or a plurality of global and local feature fusion units run in parallel; and / or the convolutional kernel sizes of every two global and local feature fusion units are different or the arrangement orders of the dilated convolutional layer and the fully convolutional layer are different.

[0011] According to some embodiments, the initial segmentation model is trained using a target organ segmentation annotation dataset to obtain a target organ segmentation prediction model, including: dividing the target organ segmentation annotation dataset into a training set, a validation set, and a test set according to a preset ratio; using a preset data augmentation algorithm to perform data augmentation on each target organ image sample in the training set, and preprocessing each target organ image sample in the data-augmented training set to obtain a target training set; based on the target training set, training the initial segmentation model with the goal of minimizing a preset loss function to obtain a training result; validating the training result based on the validation set and the test set, and outputting the training result as the target organ segmentation prediction model when the end condition is satisfied.

[0012] According to some embodiments, inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain a target organ simulated segmentation result, including: inputting the image data to be segmented of the target organ into the encoder of the target organ segmentation prediction model to output a target weighted multi-scale feature; inputting the target weighted multi-scale feature into the decoder of the target organ segmentation prediction model to output a target feature map; inputting the target feature map into the output layer of the target organ segmentation prediction model to output a segmentation probability map; using a preset threshold to perform binary processing on the segmentation probability map to obtain the target organ simulated segmentation result.

[0013] According to some embodiments, the image data to be segmented of the target organ is input into the encoder of the target organ segmentation prediction model, and target weighted multi-scale features are output, including: S1: Input the image data to be segmented of the target organ into the input layer as the current image to be processed; S2: Input the current image to be processed into multiple global and local feature fusion units of the current encoding layer of the encoder to obtain multiple weighted features; S3: Concatenate the multiple weighted features to obtain a concatenated feature map; S4: Discard features of the concatenated feature map according to a preset ratio to obtain the fused feature map of the current encoding layer; S5: Update the current encoding layer, and use the fused feature map as the current image to be processed, and repeat steps S2 - S5 until all encoding layers are traversed, so as to output the current fused feature map as the target weighted multi-scale features.

[0014] According to some embodiments, an initial segmentation model is constructed based on the pre-constructed global and local feature fusion unit and the U-Net network architecture, further including: using the multi-scale global and local feature fusion module as the decoding layer, and connecting multiple decoding layers based on the U-Net network architecture to form a decoder.

[0015] According to some embodiments, before inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain the simulated segmentation result of the target organ, it further includes: preprocessing the image data to be segmented.

[0016] According to one aspect of the present application, a segmentation device for organ images includes: a data construction module for constructing a target organ segmentation annotation data set, where the target organ segmentation annotation data set includes multiple groups of target organ image samples with segmentation annotations; a model construction module for constructing an initial segmentation model based on the pre-constructed global and local feature fusion unit and the U-Net network architecture; a model training module for training the initial segmentation model using the target organ segmentation annotation data set to obtain a target organ segmentation prediction model; a model inference module for inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain the simulated segmentation result of the target organ.

[0017] According to one aspect of the present application, an electronic device is provided, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.

[0018] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method as described above is implemented.

[0019] Beneficial effects

[0020] Through the above embodiments provided by the present application, the beneficial effects of the present application are as follows: A global and local feature fusion unit is introduced into the U-Net network structure to construct a target organ segmentation prediction model. The batch normalization layer of the global and local feature fusion unit can alleviate problems such as gradient disappearance or explosion in the training of deep networks, stabilize the training process. The fully convolutional layer and dilated convolutional layer of the global and local feature fusion unit can achieve the complementarity of global and local information, ensuring the accuracy of the model. The attention weighting block of the global and local feature fusion unit can suppress features irrelevant to the target organ while enhancing the key target regions, reducing the overall number of model parameters. Thus, the model reduces the number of model parameters on the premise of ensuring high accuracy, solves the problems of poor robustness and low accuracy in manual delineation in the prior art, and realizes the segmentation of organ images with good robustness and high accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit the present application.

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without exceeding the scope claimed by the present application.

[0023] Figure 1 It is a flowchart of the method for segmenting organ images provided by the embodiments of the present application;

[0024] Figure 2 It is a flowchart of constructing an initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture provided by the embodiments of the present application;

[0025] Figure 3 It is a schematic structural diagram of the multi-scale global and local feature fusion module provided by the embodiments of the present application;

[0026] Figure 4 It is a flowchart of training the initial segmentation model with a target organ segmentation annotation dataset to obtain a target organ segmentation prediction model provided by the embodiments of the present application;

[0027] Figure 5 It is a flowchart of inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain a simulated segmentation result of the target organ provided by the embodiments of the present application;

[0028] Figure 6A flowchart for inputting the image data of a target organ to be segmented into an encoder of a target organ segmentation prediction model and outputting target weighted multi-scale features provided by an embodiment of the present application;

[0029] Figure 7 A block diagram of a segmentation device for organ images provided by an embodiment of the present application;

[0030] Figure 8 A schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0031] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.

[0032] In addition, the described features, structures or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, many specific details are provided to give a full understanding of the embodiments of the present application. However, those skilled in the art will realize that the technical solutions of the present application can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known methods, devices, implementations or operations are not shown or described in detail to avoid obscuring various aspects of the present application.

[0033] The block diagrams shown in the accompanying drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0034] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0035] It should be understood that although terms such as first, second, and third may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Therefore, the first component described below can be called the second component without departing from the teachings of the concept of the present application. As used herein, the term "and / or" includes any one of the associated listed items and all combinations of one or more of them.

[0036] For the specific implementation method, reference can be made to the following embodiments.

[0037] Figure 1 The following is a flowchart of the organ image segmentation method provided by the embodiments of the present application. As Figure 1 shown, the method includes steps S110 - S140.

[0038] In step S110, a target organ segmentation annotation dataset is constructed, where the target organ segmentation annotation dataset includes multiple groups of target organ image samples with segmentation annotations.

[0039] The target organ is the organ that needs to be segmented and outlined in the image. For example, organs with low contrast, variable morphology, and easy to be confused with blood vessels, including head and neck OARs (such as parotid gland, temporal lobe, spinal cord), thyroid gland, etc.

[0040] This application is particularly applicable to the segmentation of images of organs at risk (OARs). At this time, the target organ is the corresponding organ at risk.

[0041] According to the exemplary embodiment, the target organ is the thyroid gland.

[0042] The image type of the target organ image sample can be a CT image or an MRI image, and this application does not limit this. This application takes the thyroid gland as the target organ as an example for illustration, but this does not represent a limitation to this application.

[0043] It should be emphasized that when the target organ is the thyroid gland, the target organ image samples in the target organ segmentation annotation dataset at least include the image of the thyroid gland after segmentation annotation, and there is no limitation on whether there are other organs in the image.

[0044] Further, in order to improve the accuracy of the model, in some embodiments, the target organ image samples in the target organ segmentation annotation dataset are set to be images that only include the thyroid gland after segmentation annotation and do not contain other organs.

[0045] Based on the above embodiments, the radiotherapy positioning CTs of 80 patients and the corresponding thyroid gland delineation annotations are selected as 80 groups of target organ image samples, so as to construct the target organ segmentation annotation dataset.

[0046] In step S120, an initial segmentation model is constructed based on the pre - constructed global and local feature fusion unit and U - Net network architecture.

[0047] Based on the original U-Net encoder-decoder framework, this application adds a Global and Local Attention (GLA) unit to form the base model of the target organ segmentation prediction model, denoted as the initial segmentation model.

[0048] According to the exemplary embodiment, a GLA unit is added after each layer of convolutional operation during the encoding stage.

[0049] In step S130, the initial segmentation model is trained using the target organ segmentation annotation dataset to obtain the target organ segmentation prediction model.

[0050] This application does not limit the specific algorithm for model training, and a suitable training algorithm can be selected according to the actual situation.

[0051] According to the exemplary embodiment, the initial segmentation model is trained using the Adam or SGD optimizer, and the initial learning rate is, for example, 1e-4.

[0052] In step S140, the image data of the target organ to be segmented is input into the target organ segmentation prediction model to obtain the simulated segmentation result of the target organ.

[0053] During the application process, the image data of the target organ to be segmented is input into the trained target organ segmentation prediction model, and the simulated segmentation result of the target organ can be obtained.

[0054] The image type of the image data to be segmented can be a CT image or an MRI image, and this application does not limit this.

[0055] Furthermore, in order to improve the accuracy, the image data to be segmented can be preprocessed before being input into the target organ segmentation prediction model.

[0056] The organ image segmentation method provided by this application can be integrated into the clinical radiotherapy planning system, and then the segmentation of the target organ image can be automatically completed in batches directly in the server environment.

[0057] This application constructs a target organ segmentation prediction model by introducing a global and local feature fusion unit into the U-Net network structure, reducing the number of model parameters while ensuring high accuracy, and realizing the segmentation of organ images with good robustness and high accuracy.

[0058] According to some embodiments, the global and local feature fusion unit includes a dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a fully convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence, or a fully convolutional layer, a first attention weighting block, a first batch normalization layer, a dilated convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence.

[0059] The global and local feature fusion unit mainly consists of three parts: Batch Normalization (BN) layer (batch normalization layer), convolutional path, and attention mechanism.

[0060] Among them, a BN layer is introduced after each convolutional operation to alleviate problems such as vanishing or exploding gradients in the training of deep networks and stabilize the training process.

[0061] Among them, the convolutional path includes "ordinary convolution" and "dilated convolution" with the same convolutional kernel size, as well as the fusion of results, to achieve the complementarity of global and local information.

[0062] Among them, the attention mechanism is implemented through channel squeeze and spatial excitation (sSE) or spatial squeeze and channel excitation (cSE) blocks, denoted as attention weighting blocks, which are used to suppress features irrelevant to the target organ and enhance key target regions. This attention mechanism is only used in the downsampling stage (encoding stage) to reduce the overall number of model parameters.

[0063] In summary, a global and local feature fusion unit includes a dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a fully convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence.

[0064] A global and local feature fusion unit can also be a fully convolutional layer, a first attention weighting block, a first batch normalization layer, a dilated convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence.

[0065] It should be explained that the first and second are only used to distinguish different layers and do not serve as limitations on order or priority.

[0066] According to some embodiments, referring to Figure 2 , in step S120, based on the pre-constructed global and local feature fusion unit and U-Net network architecture, an initial segmentation model is constructed, which can be specifically implemented through steps S210 - S230.

[0067] In step S210, multiple global and local feature fusion units are concatenated to construct a multi-scale global and local feature fusion module.

[0068] When concatenating multiple global and local feature fusion units, it should be noted that the convolutional kernel sizes of the ordinary convolution and dilated convolution of multiple global and local feature fusion units can be the same or different.

[0069] This application also does not limit the number of global and local feature fusion units.

[0070] In some embodiments, such as Figure 3 shown, a multi-scale global and local feature fusion module is constituted by simultaneously introducing multiple convolution kernel sizes (e.g., 3×3, 5×5, and 7×7) and a global and local feature fusion unit of a dilated convolution path, constituting a multi-scale convolution path, obtaining features under different receptive fields to further enhance the complementarity of global and local information, thereby improving the model accuracy.

[0071] In step S220, the multi-scale global and local feature fusion module is used as an encoding layer, and multiple encoding layers are connected based on the U-Net network architecture to form an encoder.

[0072] It can be understood that the U-Net network architecture consists of a contracting path (encoder) and an expanding path (decoder) to form a symmetric U-shaped structure, and multi-level feature fusion is achieved through skip connections.

[0073] In this application, the encoding layer of the encoder in the original U-Net network architecture is replaced with a multi-scale global and local feature fusion module to form a new encoder.

[0074] In this application, multi-scale convolution and dilated convolution are simultaneously incorporated in feature extraction, combined with channel / space attention, to suppress areas prone to confusion such as blood vessels and soft tissues, which can effectively reduce over-segmentation or under-segmentation of the contour edges of the target organ.

[0075] In step S230, based on the U-Net network architecture, an input layer, an encoder, a pre-constructed decoder, and an output layer are connected to obtain an initial segmentation model.

[0076] According to the exemplary embodiment, the input layer, decoder, and output layer in the original U-Net network architecture are used and connected to the new encoder constructed in step S220 to obtain an initial segmentation model.

[0077] Furthermore, other decoders can also be used for the decoder. For example, in some embodiments, a BN layer is added to the decoding layer of the decoder in the original U-Net network architecture, and the BN layer combines with the residual structure to reduce the risk of gradient vanishing and improve training stability.

[0078] According to some embodiments, in step S120, based on the pre-constructed global and local feature fusion unit and the U-Net network architecture, constructing an initial segmentation model further includes step S121.

[0079] In step S121, the global and local feature fusion unit is residually connected to the input layer.

[0080] Each global and local feature fusion unit is connected to the input layer through a residual block for skip connection, that is, the feature map output by the global and local feature fusion unit is concatenated with the input feature map to retain detailed information and avoid losing important spatial information. The specific connection method of the skip connection can refer to Figure 3 and the residual connection is the short connection in the figure.

[0081] According to some embodiments, the method further includes one or more of step S240, step S250, and step S260.

[0082] In step S240, the dilated convolutional layer and the fully convolutional layer of each global and local feature fusion unit run in parallel.

[0083] It should be noted that for the fully convolutional layer and the dilated convolutional layer with the same convolutional kernel size in each global and local feature fusion unit, during the running process, ordinary convolution and dilated convolution are extracted in parallel, and then the results are fused.

[0084] In step S250, multiple global and local feature fusion units run in parallel.

[0085] To improve efficiency, multiple global and local feature fusion units in the multi-scale global and local feature fusion module run in parallel. Each global and local feature fusion unit is connected to the input layer, receives the image data input by the input layer, and performs parallel processing simultaneously.

[0086] The multiple parallel global and local feature fusion units can refer to Figure 3 .

[0087] In step S260, the convolutional kernel sizes of every two global and local feature fusion units are different or the arrangement orders of the dilated convolutional layer and the fully convolutional layer are different.

[0088] Take Figure 3 as an example for illustration. To obtain features with different receptive fields, a 3×3 dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a 3×3 fully convolutional layer, a second attention weighting block, and a second batch normalization layer form the first global and local feature fusion unit; a 3×3 fully convolutional layer, a first attention weighting block, a first batch normalization layer, a 3×3 dilated convolutional layer, a second attention weighting block, and a second batch normalization layer form the second global and local feature fusion unit; a 5×5 fully convolutional layer, a first attention weighting block, a first batch normalization layer, a 5×5 dilated convolutional layer, a second attention weighting block, and a second batch normalization layer form the third global and local feature fusion unit.

[0089] The arrangement order of the dilated convolutional layer and the fully convolutional layer of the first global and local feature fusion unit and the second global and local feature fusion unit is different. The convolutional kernel sizes of the second global and local feature fusion unit and the third global and local feature fusion unit are different.

[0090] According to some embodiments, referring to Figure 4 , in step S130, the initial segmentation model is trained using the target organ segmentation annotation dataset to obtain the target organ segmentation prediction model, which can be specifically implemented through steps S410 - S440.

[0091] In step S410, the target organ segmentation annotation dataset is divided into a training set, a validation set, and a test set according to a preset ratio.

[0092] According to the exemplary embodiment, the target organ segmentation annotation dataset is divided into a training set, a validation set, and a test set in a ratio of 6:1:1.

[0093] In step S420, a preset data augmentation algorithm is used to perform data augmentation on each target organ image sample in the training set, and each target organ image sample in the training set after data augmentation is preprocessed to obtain the target training set.

[0094] The preset data augmentation algorithm includes, but is not limited to, data augmentation methods such as rotation, flipping, scaling, and shearing.

[0095] The preset data augmentation algorithm is used to process each target organ image sample in the training set, thereby expanding the dataset and improving the generalization ability of the model.

[0096] The preprocessing includes, but is not limited to, contrast enhancement methods and / or normalization methods such as HU value (Hounsfield unit) conversion and adaptive histogram equalization.

[0097] Performing secondary processing on the training set after data augmentation in a preprocessing manner can improve the image contrast and perform normalization processing to highlight the regional features of the target organ.

[0098] In step S430, based on the target training set, with the goal of minimizing the preset loss function, the initial segmentation model is trained to obtain the training result.

[0099] According to the exemplary embodiment, the preset loss function is Dice Loss. Dice Loss can maintain sensitivity to small targets under the condition of unbalanced positive and negative samples and improve the accuracy of the model for the target organ region.

[0100] The Dice Loss formula is as follows:

[0101] ;

[0102] Among them, X represents the segmentation annotation, while Y represents the prediction of the model. is the smoothing coefficient.

[0103] According to the exemplary embodiment, Take 1e-5.

[0104] In step S440, based on the validation set and the test set, the training result is verified. When the end condition is met, the training result is output as the target organ segmentation prediction model.

[0105] During the actual training process, the training batch can be set according to the scale of the training set. For each training batch, the validation set is used for verification, and the best model weights are retained.

[0106] After the final training is completed, the test set is used to test the training result, and the training result that meets the end condition is output as the target organ segmentation prediction model.

[0107] The end condition can be that a certain metric reaches a first preset value, or the number of training epochs reaches a second preset value. This application does not limit this.

[0108] According to some embodiments, referring to Figure 5 , in step S140, the image data of the target organ to be segmented is input into the target organ segmentation prediction model to obtain the target organ simulated segmentation result, which can be specifically implemented through steps S510 - S540.

[0109] In step S510, the image data of the target organ to be segmented is input into the encoder of the target organ segmentation prediction model, and the target weighted multi-scale features are output.

[0110] During the actual application process, the image data of the target organ to be segmented is input through the input layer into the encoder of the target organ segmentation prediction model for encoding.

[0111] During the encoding process, each layer of the network performs multi-scale feature extraction and attention weighting through the multi-scale global and local feature fusion module, and the target weighted multi-scale features are output.

[0112] In step S520, the target weighted multi-scale features are input into the decoder of the target organ segmentation prediction model, and the target feature map is output.

[0113] The target weighted multi-scale features output by the encoder are input into the decoder of the target organ segmentation prediction model for decoding.

[0114] During the decoding process, the residual structure and BN are adopted to reduce the risk of gradient disappearance and improve the training stability, and the target feature map is output.

[0115] In step S530, the target feature map is input into the output layer of the target organ segmentation prediction model to output a segmentation probability map.

[0116] In the output layer of the target organ segmentation prediction model, the sigmoid activation function is used so that the output value is between 0 and 1, that is, an image is generated where the value of each pixel is between 0 and 1, denoted as the segmentation probability map. This probability map represents the possibility of each pixel belonging to the foreground.

[0117] In step S540, the segmentation probability map is binarized using a preset threshold to obtain the target organ simulated segmentation result.

[0118] According to the exemplary embodiment, the preset threshold is set to 0.5. That is, if the probability value of a pixel is greater than 0.5, the pixel is classified as the foreground (i.e., the target organ), and if the probability value of the pixel is less than or equal to 0.5, the pixel is classified as the background.

[0119] Pixels with probability values greater than the preset threshold are taken as the segmented target organ image to obtain the target organ simulated segmentation result.

[0120] According to some embodiments, referring to Figure 6 , in step S510, the image data of the target organ to be segmented is input into the encoder of the target organ segmentation prediction model to output the target weighted multi-scale features, which can be specifically implemented through steps S1 - S5.

[0121] S1: The image data of the target organ to be segmented is input into the input layer as the current image to be processed.

[0122] The input layer is used to receive the original unprocessed image and use it as the current image to be processed.

[0123] S2: The current image to be processed is input into multiple global and local feature fusion units of the current encoding layer of the encoder to obtain multiple weighted features.

[0124] In the first round of loop, the first encoding layer is used as the current encoding layer, and the current image to be processed is input into multiple global and local feature fusion units of the current encoding layer for multi-scale feature extraction and attention weighting to obtain multiple weighted features.

[0125] S3: The multiple weighted features are concatenated to obtain a concatenated feature map.

[0126] S4: The concatenated feature map is subjected to feature discard according to a preset ratio to obtain the fusion feature map of the current encoding layer.

[0127] According to an exemplary embodiment, the preset ratio is 0.3. It should be noted that the core purpose of feature discarding according to the preset ratio is to enhance the generalization ability of the model and prevent overfitting.

[0128] S5: Update the current encoding layer, and use the fused feature map as the current image to be processed, and repeat steps S2 - S5 until all encoding layers are traversed, so as to output the current fused feature map as the target weighted multi-scale feature.

[0129] Take the next encoding layer as the current encoding layer, perform encoding processing layer by layer in order, and use the current fused feature map output by the last encoding layer as the target weighted multi-scale feature.

[0130] According to some embodiments, in step S120, based on a pre-constructed global and local feature fusion unit and U-Net network architecture, an initial segmentation model is constructed, and step S122 is further included.

[0131] In step S122, use the multi-scale global and local feature fusion module as the decoding layer, and connect multiple decoding layers based on the U-Net network architecture to form a decoder.

[0132] Based on the above embodiments, the essence of using the multi-scale global and local feature fusion module as the encoding layer is to change the two convolutions in each layer of the original U-Net encoder into a multi-scale global and local feature fusion module.

[0133] Referring to the above modification, change the two convolutions in each layer of the original U-Net decoder into a multi-scale global and local feature fusion module.

[0134] According to some embodiments, before step S140, step S139 is further included.

[0135] In step S139, preprocess the image data to be segmented.

[0136] The preprocessing of the image data to be segmented is similar to the preprocessing method of the target organ image sample, and the present application will not elaborate here.

[0137] Furthermore, after step S140, step S150 can be further included.

[0138] In step S150, the annotator is used to finely adjust a few suspicious boundaries of the simulated segmentation result of the target organ.

[0139] The present application reduces the manual participation, reduces the labor cost, while reducing the delineation variability and inconsistency caused by differences in level and experience, and greatly shortens the OAR delineation time.

[0140] To further illustrate the organ image segmentation method provided in this application, a specific embodiment is given.

[0141] In one embodiment, the localization CTs of 80 head and neck or breast cancer patients and their thyroid gland delineation annotations are obtained from a clinical database, and the training set, validation set, and test set are divided according to a ratio of 6:1:1, and data augmentation is performed. During training, an Adam or SGD optimizer is used, with an initial learning rate such as 1e-4, and it is dynamically adjusted during training. According to the scale of the training set, generally about 100 - 300 Epochs (batches) are set; after each Epoch, the best model weights are monitored and saved on the validation set to obtain a target organ segmentation prediction model.

[0142] The preprocessed patient CT is input into the trained target organ segmentation prediction model to obtain a probability map output; then, thresholding (such as 0.5) is used for binarization processing to obtain the final segmentation result.

[0143] The experimental results show that the organ image segmentation method provided in this application is superior to traditional U-Net, Attention U-Net, HR-Net, and other improved models in terms of indicators such as Dice coefficient (DSC), Jaccard coefficient (JSC), positive predictive value (PPV), sensitivity (SE), Hausdorff distance (HD), relative volume difference (RVD), and voxel overlap error (VOE), and can achieve segmentation accuracy comparable to that of experienced annotators. Compared with some improved network structures or traditional U-Net models, this application reduces the network parameter quantity and FLOPS while ensuring high accuracy, which is conducive to efficient deployment in resource-constrained environments (such as ordinary hospital servers or mobile devices). The automatic segmentation results can reduce the workload of physicians' repeated delineation and improve the efficiency of radiotherapy plan formulation.

[0144] The device embodiments of this application are described below, which can be used to execute the method embodiments of this application. For details not disclosed in the device embodiments of this application, reference can be made to the method embodiments of this application.

[0145] Figure 7 The block diagram of an organ image segmentation device according to an exemplary embodiment is shown.

[0146] Figure 7 The device shown can execute the organ image segmentation method according to the embodiments of this application described above.

[0147] As Figure 7 shown, the organ image segmentation device may include:

[0148] See Figure 7With reference to the previous description, the data construction module 710 is used to construct a target organ segmentation annotation dataset, where the target organ segmentation annotation dataset includes multiple groups of target organ image samples with segmentation annotations.

[0149] The model construction module 720 is used to construct an initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture.

[0150] The model training module 730 is used to train the initial segmentation model using the target organ segmentation annotation dataset to obtain a target organ segmentation prediction model.

[0151] The model inference module 740 is used to input the image data of the target organ to be segmented into the target organ segmentation prediction model to obtain a simulated segmentation result of the target organ.

[0152] The device performs functions similar to the method provided previously. For other functions, please refer to the previous description and will not be elaborated here.

[0153] Figure 8 An electronic device according to an exemplary embodiment of the present application is shown. The following refers to Figure 8 to describe the electronic device 800 according to this embodiment of the present application. Figure 8 The electronic device 800 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.

[0154] As Figure 8 shown, the electronic device 800 is presented in the form of a general computing device. The components of the electronic device 800 may include but are not limited to: at least one processing unit 810, at least one storage unit 820, a display unit 840, etc.

[0155] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 810, so that the processing unit 810 executes the methods according to various exemplary embodiments of the present application described in this specification. For example, the processing unit 810 can execute the method as described above.

[0156] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.

[0157] The storage unit 820 may further include a program / utility 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0158] The electronic device 800 can also communicate with one or more external devices 300 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 800, and / or communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 850. Moreover, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0159] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. The technical solutions according to the embodiments of the present application can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiments of the present application.

[0160] The software product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0161] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0162] The program code for performing the operations of this application may be written in any combination of one or more programming languages. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0163] The above computer-readable medium carries one or more programs, which, when executed by such a device, cause the computer-readable medium to implement the foregoing functions.

[0164] Those skilled in the art can understand that the above-mentioned modules may be distributed in the device according to the description of the embodiments, or may be correspondingly changed and distributed in one or more devices that are only different from this embodiment. The modules of the above embodiments may be combined into one module, or may be further split into multiple sub-modules.

[0165] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, server, mobile terminal, or network device, etc.) to execute the method according to the embodiments of this application.

[0166] The exemplary embodiments of the present application have been specifically shown and described above. It should be understood that the present application is not limited to the detailed structures, arrangements or implementation methods described herein; on the contrary, the present application is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A method for segmenting an organ image, characterized in that, Including: Constructing a target organ segmentation annotation dataset, where the target organ segmentation annotation dataset includes multiple groups of target organ image samples with segmentation annotations; Constructing an initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture; Training the initial segmentation model using the target organ segmentation annotation dataset to obtain a target organ segmentation prediction model; Inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain a simulated segmentation result of the target organ; The global and local feature fusion unit includes a dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a fully convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence, or a fully convolutional layer, a first attention weighting block, a first batch normalization layer, a dilated convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence; The constructing of the initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture includes: Stitching multiple of the global and local feature fusion units to construct a multi-scale global and local feature fusion module; Using the multi-scale global and local feature fusion module as an encoding layer, and connecting multiple of the encoding layers based on the U-Net network architecture to form an encoder; Based on the U-Net network architecture, connecting an input layer, the encoder, a pre-constructed decoder, and an output layer to obtain the initial segmentation model.

2. The method according to claim 1, wherein The constructing of the initial segmentation model based on a pre-constructed global and local feature fusion unit and a U-Net network architecture further includes: Residually connecting the global and local feature fusion unit and the input layer.

3. The method according to claim 1, wherein The method further includes: The dilated convolutional layer and the fully convolutional layer of each global and local feature fusion unit run in parallel; and / or Multiple of the global and local feature fusion units run in parallel; and / or The convolutional kernel sizes of every two of the global and local feature fusion units are different or the arrangement order of the dilated convolutional layer and the fully convolutional layer is different.

4. The method according to claim 1, wherein The training of the initial segmentation model using the target organ segmentation annotation dataset to obtain a target organ segmentation prediction model includes: Dividing the target organ segmentation annotation dataset into a training set, a validation set, and a test set according to a preset ratio; Using a preset data augmentation algorithm to perform data augmentation on each target organ image sample in the training set, and preprocessing each target organ image sample in the training set after data augmentation to obtain a target training set; Based on the target training set, training the initial segmentation model with the goal of minimizing a preset loss function to obtain a training result; Based on the validation set and the test set, validating the training result, and outputting the training result as a target organ segmentation prediction model when the end condition is met.

5. The method according to claim 1, wherein The inputting of the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain a simulated segmentation result of the target organ includes: Input the image data to be segmented of the target organ into the encoder of the target organ segmentation prediction model, and output the target weighted multi-scale features; Input the target weighted multi-scale features into the decoder of the target organ segmentation prediction model, and output the target feature map; Input the target feature map into the output layer of the target organ segmentation prediction model, and output the segmentation probability map; Use a preset threshold to perform binarization processing on the segmentation probability map to obtain the simulated segmentation result of the target organ.

6. The method according to claim 5, characterized in that, The step of inputting the image data to be segmented of the target organ into the encoder of the target organ segmentation prediction model and outputting the target weighted multi-scale features includes: S1: Input the image data to be segmented of the target organ into the input layer as the current image to be processed; S2: Input the current image to be processed into multiple global and local feature fusion units of the current encoding layer of the encoder to obtain multiple weighted features; S3: Concatenate the multiple weighted features to obtain a concatenated feature map; S4: Discard features of the concatenated feature map according to a preset ratio to obtain the fused feature map of the current encoding layer; S5: Update the current encoding layer, and use the fused feature map as the current image to be processed, and repeat steps S2 - S5 until all encoding layers are traversed, so as to output the current fused feature map as the target weighted multi-scale features.

7. The method according to claim 1, characterized in that, Based on the pre-constructed global and local feature fusion unit and the U-Net network architecture, constructing the initial segmentation model further includes: Use the multi-scale global and local feature fusion module as the decoding layer, and connect multiple decoding layers based on the U-Net network architecture to form a decoder.

8. The method according to claim 1, characterized in that, Before inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain the simulated segmentation result of the target organ, it further includes: Preprocess the image data to be segmented.

9. An organ image segmentation device, characterized in that, It includes: A data construction module for constructing a target organ segmentation annotation data set, where the target organ segmentation annotation data set includes multiple groups of target organ image samples with segmentation annotations; A model construction module for constructing an initial segmentation model based on the pre-constructed global and local feature fusion unit and the U-Net network architecture; A model training module for training the initial segmentation model using the target organ segmentation annotation data set to obtain a target organ segmentation prediction model; A model inference module for inputting the image data to be segmented of the target organ into the target organ segmentation prediction model to obtain the simulated segmentation result of the target organ; The global and local feature fusion unit includes a dilated convolutional layer, a first attention weighting block, a first batch normalization layer, a fully convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence, or a fully convolutional layer, a first attention weighting block, a first batch normalization layer, a dilated convolutional layer, a second attention weighting block, and a second batch normalization layer connected in sequence; Based on the pre-constructed global and local feature fusion unit and the U-Net network architecture, constructing the initial segmentation model includes: Splice multiple of the global and local feature fusion units to construct a multi-scale global and local feature fusion module; Use the multi-scale global and local feature fusion module as an encoding layer, and connect multiple of the encoding layers based on the U-Net network architecture to form an encoder; Based on the U-Net network architecture, connect an input layer, the encoder, a pre-constructed decoder, and an output layer to obtain the initial segmentation model.

10. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-8.

11. A computer-readable storage medium having a computer program or instructions stored thereon, characterized in that, When the computer program or instruction is executed by a processor, the method according to any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Abdominal CT image multi-organ automatic segmentation method

    CN117994270A

  • Abdomen multi-organ image segmentation method fusing multi-scale features

    CN119205824A