Medical image segmentation method and device, computer equipment and storage medium
By introducing multiple segmentation models, feature interaction models and fusion models into the medical image segmentation method, the shortcomings of existing methods in capturing remote dependencies and detail perception capabilities are solved, and the effect of medical image segmentation is significantly improved.
Patent Information
- Application Number
- CN202410836830.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-26
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-06-26
AI Technical Summary
The existing medical image segmentation methods have insufficient ability to capture remote dependencies and detail perception, especially when processing medical images, the generalization and representation capabilities of the model are insufficient, resulting in poor segmentation effect.
A medical image segmentation method is proposed, by acquiring medical images and performing image segmentation based on a trained medical image segmentation model. The model includes multiple segmentation models, feature interaction models and fusion models. The feature interaction model allows each segmentation model to learn from each other, thereby improving the segmentation effect.
It significantly improves the effect of medical image segmentation, enhances the model's global modeling ability and detail perception ability, and improves the model's generalization ability and representation ability.
Smart Images

Figure CN120088273A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image segmentation methods, and particularly to a medical image segmentation method, device, computer device, and storage medium. Background Technique
[0002] Medical imaging technology is a very important technical means in the medical field. Through different physical principles and devices, it can obtain information on the internal structure and function of the human body, helping doctors diagnose diseases and guide treatments. Common medical imaging technologies include CT, MRI, PET, etc. These technologies are widely used due to their painless and non-invasive characteristics and the ability to provide high-quality image information. However, extracting useful information from these high-quality images is a challenge, which highly depends on doctors' clinical experience. Therefore, advanced image analysis tools are needed to process these image data, which can help doctors better understand diseases, formulate more scientific treatment plans, and improve the medical level and the quality of life of patients.
[0003] In the field of medical image segmentation, multiple deep learning-based methods have been proposed to extract regions and structures of interest in medical images for treatment detection and disease diagnosis. Since the pioneering U-Net network was proposed, CNN-based networks have quickly achieved state-of-the-art results in various 2D and 3D medical image segmentation tasks due to their strong detail perception ability, strong adaptability, and ease of training. These networks extract data features through convolutional layers and learn the mapping relationship between the original image and the segmentation result from a large amount of training data. To make up for the shortcoming that convolutional layers cannot capture long-range dependencies, Transformer-based methods convert the image segmentation problem into a sequence prediction problem and effectively overcome the shortcoming that convolutional layers cannot capture long-range dependencies by introducing Transformer modules. Finally, with the powerful global modeling ability of Transformer, the segmentation performance of the model is effectively improved, especially in dealing with long-range dependencies and images of different scales. The SAM model is a large segmentation base model that has been pre-trained with more than 1 billion masks on 11 million natural images. Thanks to the large amount of training data and the general model architecture, SAM shows amazing zero-shot performance in various natural image segmentation tasks. The SAM model consists of three parts: an image encoder, a prompt encoder, and a mask generator. The prompt encoder receives various prompt information such as points, boxes, and texts to enhance the model's feature learning ability. Compared with traditional segmentation networks, the SAM model has the advantages of strong generalization, strong robustness, and strong feature representation ability.
[0004] Although many algorithms for medical image segmentation have been developed, there are still many deficiencies in existing methods. CNN-based segmentation methods are limited by the structure of convolutional kernels and cannot capture long-range dependencies well, which is not conducive to medical image segmentation tasks. Although Transformer-based methods can effectively improve this phenomenon, Transformer has the disadvantages of slow convergence speed and weak detail perception ability. Therefore, a large amount of data is often required to make the model converge to an ideal state. However, this is a challenge for scarce medical data. In addition, since existing segmentation methods are trained based on very limited data, the model cannot learn powerful representation capabilities, which leads to poor performance of the model when facing complex data and also causes a significant decline in the generalization ability of the model. The SAM model shows impressive performance in natural image segmentation tasks. However, due to the significant differences between medical images and natural images, directly applying the SAM model to medical image tasks is not ideal, or even poor. In addition, as a two-dimensional model, SAM cannot utilize the three-dimensional information in medical images, which is not conducive to medical image segmentation tasks. Summary of the Invention
[0005] Based on this, in view of the technical problem of the poor medical image segmentation effect of the existing technology, it is necessary to propose a medical image segmentation method, device, computer device and storage medium.
[0006] In a first aspect, a medical image segmentation method is provided. The method includes:
[0007] Obtain a medical image;
[0008] Perform image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Wherein, the medical image segmentation model includes each segmentation model, the feature interaction model corresponding to each segmentation model, and a fusion model.
[0009] In a second aspect, a medical image segmentation device is provided. The device includes:
[0010] An acquisition module, configured to acquire a medical image;
[0011] An image segmentation module, configured to perform image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Wherein, the medical image segmentation model includes each segmentation model, the feature interaction model corresponding to each segmentation model, and a fusion model.
[0012] In a third aspect, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above-mentioned medical image segmentation method are implemented.
[0013] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned medical image segmentation method are implemented.
[0014] The medical image segmentation method proposed by the present invention obtains a medical image, and then performs image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Among them, the medical image segmentation model includes each segmentation model, the respective feature interaction models corresponding to each segmentation model, and a fusion model, and can enable the segmentation models to learn from each other through the feature interaction models, thereby significantly improving the effect of medical image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Among them:
[0017] Figure 1 It is an application environment diagram of the medical image segmentation method in an embodiment;
[0018] Figure 2 It is a flowchart of the medical image segmentation method in an embodiment;
[0019] Figure 3 It is the model structure of the medical image segmentation model of the medical image segmentation method in an embodiment;
[0020] Figure 4 It is the model structure of the feature interaction model of the medical image segmentation method in an embodiment;
[0021] Figure 5 It is the model structure of the fusion model of the medical image segmentation method in an embodiment;
[0022] Figure 6 It is a comparison result of the medical image segmentation method in an embodiment;
[0023] Figure 7Another comparison result of the medical image segmentation method in one embodiment;
[0024] Figure 8 A visualization result of the medical image segmentation method in one embodiment;
[0025] Figure 9 Another visualization result of the medical image segmentation method in one embodiment;
[0026] Figure 10 The structural block diagram of the medical image segmentation method device in one embodiment;
[0027] Figure 11 The structural block diagram of the computer device in one embodiment;
[0028] Figure 12 The structural block diagram of the computer device in another embodiment. Detailed implementation manners
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification and claims of this application or the above drawings are used to distinguish different objects and not to describe a specific order.
[0030] Referring to "embodiment" herein means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0032] The medical image segmentation method provided by the embodiments of the present invention can be applied in, for example Figure 1In the application environment, the client 110 communicates with the server 120 through a network. The server 120 can use the medical image segmentation method proposed in this embodiment. By obtaining a medical image, and then the server 120 performs image segmentation based on the medical image and the trained medical image segmentation model to obtain the segmentation result of the medical image. Among them, the medical image segmentation model includes each segmentation model, the feature interaction model corresponding to each segmentation model, and the fusion model, which can enable the segmentation models to learn from each other through the feature interaction model, thus significantly improving the effect of medical image segmentation. Among them, the client 110 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail through specific embodiments below.
[0033] Please refer to Figure 2 as shown in Figure 2 FIG. is a schematic flowchart of a medical image segmentation method provided by an embodiment of the present invention, including the following steps:
[0034] Step S101: Obtain a medical image;
[0035] Step S102: Perform image segmentation based on the medical image and the trained medical image segmentation model to obtain the segmentation result of the medical image. Among them, the medical image segmentation model includes each segmentation model, the feature interaction model corresponding to each segmentation model, and the fusion model.
[0036] Among them, the segmentation model can be models such as U-Net, nnU-Net, UNETR, Swin-UNETR, etc. The feature interaction model is obtained by training a model based on the attention mechanism model and convolution. The feature interaction model is used to perform feature interaction learning on each feature extracted from the medical image by the segmentation model. The fusion model can be obtained by training a model based on a convolutional neural network.
[0037] It should be noted that models such as U-Net, nnU-Net, UNETR, and Swin-UNETR have all achieved good results due to their own advantages, but they also have their own deficiencies. For example, problems such as the inability to capture long-range dependencies and weak detail perception ability also restrict the development of the models. In recent years, the powerful generalization ability and emergence ability of large models have gradually promoted the transformation of the deep learning model development paradigm. However, training large models requires huge computing resources, and at the same time, the huge amount of data required is also difficult to achieve in the field of medical images. Although the SAM model has achieved great success in the field of natural images, due to the huge differences between natural images and medical images, directly applying SAM to medical images does not yield ideal results. However, the powerful generalization ability of SAM makes it have the potential to be applied to medical image segmentation.
[0038] The model structure of the medical image segmentation model is as Figure 3 shown. By constructing a large model cluster to give full play to the advantages of the model itself, the optimal segmentation result is finally achieved. The large model cluster refers to each segmentation model and the corresponding feature interaction models of each segmentation model. In addition, for existing basic large models, rapid model migration can also be achieved through continuous interaction with the proprietary models trained on medical images within the cluster.
[0039] In one embodiment, the segmentation model includes a first segmentation model, a second segmentation model, and a third segmentation model. The feature interaction models include a first feature interaction model corresponding to the first segmentation model, a second feature interaction model corresponding to the second segmentation model, and a third feature interaction model corresponding to the third segmentation model. The step of performing image segmentation on the medical image based on the medical image and the trained medical image segmentation model to obtain the segmentation result of the medical image includes:
[0040] Step S201: Input the medical image into the first segmentation model to obtain the first feature output by the first segmentation model;
[0041] Step S202: Input the medical image into the second segmentation model to obtain the second feature output by the second segmentation model;
[0042] Step S203: Input the medical image into the third segmentation model to obtain the third feature output by the third segmentation model;
[0043] Step S204: Perform image segmentation based on the first feature, the second feature, the third feature, the first feature interaction model, the second feature interaction model, the third feature interaction model, and the fusion model to obtain the segmentation result of the medical image.
[0044] Among them, the first segmentation model can be a SAM model, the second segmentation model can be a U-Net model, and the third segmentation model can be a UNETR model. These three segmentation models respectively represent a basic large model with strong generalization ability, a CNN model with strong detail perception ability, and a Transformer model with strong global modeling ability.
[0045] As an example, SAM model: In order to enhance the utilization of the original SAM model for medical 3D information, the processed medical image is first sent to the 3D adapter to extract volume information. m is defined as the feature map and F is the output of the 3D adapter. The process can be expressed as:
[0046] F(m)=m+σ(Conv3D(Norm(m)))
[0047] Among them, σ is the activation function, Norm represents the normalization layer, and Conv3D represents the 3D convolution layer, which is used to extract 3D information. Then the output F(m) of the 3D adapter is sent to the SAM image encoder for further feature extraction and generates the encoded output o i , O i That's it.
[0048] o i =SAM(F(m))
[0049] UNTER model: The UNETR model is a medical image segmentation model designed based on Transformer. With the powerful global modeling capability of Transformer, the UNETR model can effectively capture the long-range dependencies in three-dimensional medical data. Define x as the model input feature, then the model output o R It is expressed as:
[0050] o R =UNETR(x)
[0051] U-Net model: The U-Net model can effectively capture detail information with its convolutional structure. Here it is mainly used to supplement detail features. The input feature is defined as x. Then the output of the model is o U It can be expressed as:
[0052] o U =UNET(x)
[0053] In one embodiment, the step of performing image segmentation based on the first feature, the second feature, the third feature, the first feature interaction model, the second feature interaction model, the third feature interaction model and the fusion model to obtain a segmentation result of the medical image includes:
[0054] Step S301: Determine a fourth feature based on the first feature, the second feature, the third feature, and the first feature interaction model;
[0055] Step S302: Determine a fifth feature based on the first feature, the second feature, the third feature, and the second feature interaction model;
[0056] Step S303: Determine a sixth feature based on the first feature, the second feature, the third feature, and the third feature interaction model;
[0057] Step S304: Perform image segmentation based on the fourth feature, the fifth feature, the sixth feature, and the fusion model to obtain a segmentation result of the medical image.
[0058] In this embodiment, different features can be mutually guided through the feature interaction model, and their own deficiencies can be made up by continuously interacting with other models. For example, the SAM model, as a basic large model, has strong generalization ability. However, since it is trained based on natural images and is a two-dimensional model, its direct application to medical images has poor effects. However, the UNETR model based on Transformer can effectively capture long-range dependencies in three-dimensional data, and the U-Net network based on CNN can fully focus on detailed information. Therefore, the SAM model can continuously interact with the UNETR model and the U-Net model through the first feature interaction model to improve its global modeling ability and detail perception ability. At the same time, under the guidance of the UNETR model and the U-Net model, the SAM model can more efficiently achieve the migration of medical image data. In addition, with the improvement of future hardware levels, the scale of the model cluster of this network framework can be continuously expanded.
[0059] In one embodiment, the step of determining the fourth feature based on the first feature, the second feature, the third feature, and the first feature interaction model includes:
[0060] Step S401: Obtain a first Key and a first Value based on a linear transformation of the second feature;
[0061] Step S402: Obtain a first Query based on a linear transformation of the third feature;
[0062] Step S403: Input the first Key, the first Value, and the first Query into the first attention model in the first feature interaction model to obtain a seventh feature;
[0063] Step S404: Input the seventh feature into the convolutional layer in the first feature interaction model to obtain the eighth feature output by the convolutional layer, and based on a linear transformation of the eighth feature, obtain the second Key and the second Value;
[0064] Step S405: Based on a linear transformation of the first feature, obtain the second Query;
[0065] Step S406: Input the second Key, the second Value, and the second Query into the second attention model in the first feature interaction model to obtain the fourth feature.
[0066] Among them, both the first attention model and the second attention model are models using the attention mechanism. The first Key refers to the key (K) in the attention mechanism, the first Value refers to the value (V) in the attention mechanism, and the first Query refers to the query (Q) in the attention mechanism.
[0067] As an example, the model structure of the first feature interaction model is as Figure 4 shown.
[0068] In one embodiment, the step of determining the fifth feature based on the first feature, the second feature, the third feature, and the second feature interaction model includes:
[0069] Step S501: Based on a linear transformation of the first feature, obtain the third Key and the third Value;
[0070] Step S502: Based on a linear transformation of the third feature, obtain the third Query;
[0071] Step S503: Input the third Key, the third Value, and the third Query into the third attention model in the second feature interaction model to obtain the ninth feature;
[0072] Step S504: Input the ninth feature into the convolutional layer in the second feature interaction model to obtain the tenth feature output by the convolutional layer, and based on a linear transformation of the tenth feature, obtain the fourth Key and the fourth Value;
[0073] Step S505: Based on a linear transformation of the second feature, obtain the fourth Query;
[0074] Step S506: Input the fourth Key, the fourth Value, and the fourth Query into the fourth attention model in the second feature interaction model to obtain the fifth feature.
[0075] In one embodiment, the step of determining the fifth feature based on the first feature, the second feature, the third feature, and the second feature interaction model includes:
[0076] Step S601: Obtain a fifth Key and a fifth Value by performing a linear transformation on the first feature;
[0077] Step S602: Obtain a fifth Query by performing a linear transformation on the second feature;
[0078] Step S603: Input the fifth Key, the fifth Value, and the fifth Query into the fifth attention model in the third feature interaction model to obtain an eleventh feature;
[0079] Step S604: Input the eleventh feature into the convolutional layer in the third feature interaction model to obtain a twelfth feature output by the convolutional layer, and obtain a sixth Key and a sixth Value by performing a linear transformation on the twelfth feature;
[0080] Step S605: Obtain a sixth Query by performing a linear transformation on the third feature;
[0081] Step S606: Input the sixth Key, the sixth Value, and the sixth Query into the sixth attention model in the third feature interaction model to obtain a sixth feature.
[0082] In one embodiment, the step of performing image segmentation based on the fourth feature, the fifth feature, the sixth feature, and the fusion model to obtain a segmentation result of a medical image includes:
[0083] Step S701: Concatenate the fourth feature, the fifth feature, and the sixth feature to obtain a target feature;
[0084] Step S702: Input the target feature into the 3D global average pooling layer in the fusion model to obtain a first vector, and input the first vector into the excitation layer in the fusion model to obtain a second vector;
[0085] Step S703: Input the target feature into the 3D global max pooling layer in the fusion model to obtain a third vector, and input the third vector into the excitation layer in the fusion model to obtain a fourth vector;
[0086] Step S704: Add the second vector and the fourth vector to obtain a fifth vector, and perform activation processing on the fifth vector through an activation function to obtain a probability matrix;
[0087] Step S705: Multiply the probability matrix with the target feature to obtain the segmentation result of the medical image.
[0088] As an example, the model structure of the fusion model is as Figure 5 shown.
[0089] It should be emphasized that the present invention proposes a brand-new social training mode. Based on this training framework, a large model cluster strategy is proposed. The models within the cluster are continuously optimized through interaction, and finally the optimal segmentation result is achieved. The present invention uses the SAM-based large model, UNETR, and 3DU-Net models to prove the feasibility of this social network framework. In addition, other existing models based on other basic large models such as GPT, CLIP, etc. and other traditional deep models such as U-Net, Mamba, Swin-UNETR, Swin-UNet, nnFormer, etc. can all be applied to the present invention. Taking the three-dimensional medical image segmentation task as an example, in addition, this network framework is also applicable to other various segmentation tasks, other computer vision tasks, and other fields such as NLP. The present invention uses the attention module and the convolution module to realize the interaction between models, and other methods for different feature interaction and fusion are applicable to the present invention. The present invention emphasizes building a model cluster to enrich feature learning, and other methods such as using parallel networks to fuse the output features of different networks are applicable to the present invention. For the SAM model, the present invention uses the FacT parameter fine-tuning technology to efficiently transfer the SAM model to the medical image segmentation task. There are many methods for parameter fine-tuning, such as additive, selective, reparameterization, etc. are all applicable to the present invention. The present invention uses the feature interaction module to guide the models to optimize each other and uses the melting module for feature supplementation to output the optimal segmentation result.
[0090] The present invention has the following advantages. By constructing a large model cluster through this training framework, the segmentation models within the cluster perform "model storms" through the feature interaction model, learn the optimal, most accurate, and most comprehensive feature information, and perform feature supplementation through the feature fusion module to achieve the optimal output result. The present invention redefines the large artificial intelligence model. Compared with a single general large model, the present invention pays more attention to the interaction of existing artificial intelligence models, builds a large model cluster based on existing artificial intelligence models, and at the same time provides a new idea for the development of future artificial intelligence models. In addition, with the improvement of future hardware levels, theoretically, the medical image segmentation method proposed by the present invention can infinitely expand the scale of the model cluster, break through the information cocoon of existing models, and include multi-modal and multi-dimensional information. Although this method has been verified in the medical image segmentation task, this method is also applicable to other computer vision tasks and other fields such as NLP.
[0091] As an example, the proposed medical image segmentation method of the present invention was verified on six prostate public datasets and the Sliver07 liver dataset respectively, and compared with classical traditional segmentation methods. The same VIT_L pre-trained weights were used in both MA-SAM and the medical image segmentation method. Figure 6 The comparison results of the present invention with other methods on six prostate datasets are shown. Figure 7 The comparison results of the present invention with other methods on the Sliver07 liver dataset are shown. It can be seen from the results that the proposed medical image segmentation method of the present invention has significant improvements compared with existing methods. In addition, the segmentation results of various methods were visualized, such as Figure 8 、 Figure 9 as shown. It can be seen from the visualization results that the proposed medical image segmentation method of the present invention is significantly better than existing methods in terms of both the integrity of segmentation and the detail perception of boundaries.
[0092] The medical image segmentation method proposed in this embodiment, by acquiring a medical image, and then performing image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. The medical image segmentation model includes each segmentation model, the respective feature interaction models corresponding to each segmentation model, and a fusion model, and can enable the segmentation models to learn from each other through the feature interaction models, thereby significantly improving the effect of medical image segmentation.
[0093] Please refer to Figure 10 as shown. In one embodiment, a medical image segmentation method device is provided. The device includes: an acquisition module 10 for acquiring a medical image;
[0094] An image segmentation module 20 for performing image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. The medical image segmentation model includes each segmentation model, the respective feature interaction models corresponding to each segmentation model, and a fusion model.
[0095] The image segmentation module 20 is configured to input the medical image into the first segmentation model to obtain a first feature output by the first segmentation model; input the medical image into the second segmentation model to obtain a second feature output by the second segmentation model; input the medical image into the third segmentation model to obtain a third feature output by the third segmentation model; and perform image segmentation based on the first feature, the second feature, the third feature, the first feature interaction model, the second feature interaction model, the third feature interaction model, and the fusion model to obtain a segmentation result of the medical image.
[0096] The image segmentation module 20 is configured to determine a fourth feature based on the first feature, the second feature, the third feature, and the first feature interaction model; determine a fifth feature based on the first feature, the second feature, the third feature, and the second feature interaction model; determine a sixth feature based on the first feature, the second feature, the third feature, and the third feature interaction model; and perform image segmentation based on the fourth feature, the fifth feature, the sixth feature, and the fusion model to obtain a segmentation result of the medical image.
[0097] The image segmentation module 20 is configured to: obtain a first Key and a first Value by performing a linear transformation on the second feature; obtain a first Query by performing a linear transformation on the third feature; input the first Key, the first Value, and the first Query into a first attention model in the first feature interaction model to obtain a seventh feature; input the seventh feature into a convolutional layer in the first feature interaction model to obtain an eighth feature output by the convolutional layer, and obtain a second Key and a second Value by performing a linear transformation on the eighth feature; obtain a second Query by performing a linear transformation on the first feature; and input the second Key, the second Value, and the second Query into a second attention model in the first feature interaction model to obtain a fourth feature.
[0098] The image segmentation module 20 is configured to: obtain a third Key and a third Value by performing a linear transformation on the first feature; obtain a third Query by performing a linear transformation on the third feature; input the third Key, the third Value, and the third Query into a third attention model in the second feature interaction model to obtain a ninth feature; input the ninth feature into a convolutional layer in the second feature interaction model to obtain a tenth feature output by the convolutional layer, and obtain a fourth Key and a fourth Value by performing a linear transformation on the tenth feature; obtain a fourth Query by performing a linear transformation on the second feature; and input the fourth Key, the fourth Value, and the fourth Query into a fourth attention model in the second feature interaction model to obtain a fifth feature.
[0099] The image segmentation module 20 is configured to obtain a fifth Key and a fifth Value based on a linear transformation of the first feature; obtain a fifth Query based on a linear transformation of the second feature; input the fifth Key, the fifth Value, and the fifth Query into a fifth attention model in a third feature interaction model to obtain an eleventh feature; input the eleventh feature into a convolutional layer in the third feature interaction model to obtain a twelfth feature output by the convolutional layer, and obtain a sixth Key and a sixth Value based on a linear transformation of the twelfth feature; obtain a sixth Query based on a linear transformation of the third feature; input the sixth Key, the sixth Value, and the sixth Query into a sixth attention model in the third feature interaction model to obtain a sixth feature.
[0100] The image segmentation module 20 is configured to splice the fourth feature, the fifth feature, and the sixth feature to obtain a target feature; input the target feature into a 3D global average pooling layer in a fusion model to obtain a first vector, and input the first vector into an activation layer in the fusion model to obtain a second vector; input the target feature into a 3D global max pooling layer in the fusion model to obtain a third vector, and input the third vector into an activation layer in the fusion model to obtain a fourth vector; add the second vector and the fourth vector to obtain a fifth vector, and perform an activation process on the fifth vector through an activation function to obtain a probability matrix; multiply the probability matrix by the target feature to obtain a segmentation result of the medical image.
[0101] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 11 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Wherein, the processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a medical image segmentation method.
[0102] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as Figure 12As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a medical image segmentation method.
[0103] In one embodiment, a computer device is proposed, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are realized:
[0104] Obtain a medical image;
[0105] Perform image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Among them, the medical image segmentation model includes various segmentation models, respective feature interaction models corresponding to the various segmentation models, and a fusion model.
[0106] The medical image segmentation method proposed in this embodiment, by obtaining a medical image and then performing image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Among them, the medical image segmentation model includes various segmentation models, respective feature interaction models corresponding to the various segmentation models, and a fusion model, can enable the various segmentation models to learn from each other through the feature interaction model, thus significantly improving the effect of medical image segmentation.
[0107] In one embodiment, a computer-readable storage medium is proposed. The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, the following steps are realized:
[0108] Obtain a medical image;
[0109] Perform image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. Among them, the medical image segmentation model includes various segmentation models, respective feature interaction models corresponding to the various segmentation models, and a fusion model.
[0110] The medical image segmentation method proposed in this embodiment obtains a medical image, and then performs image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image. The medical image segmentation model includes various segmentation models, respective feature interaction models corresponding to the segmentation models, and a fusion model, and can enable the segmentation models to learn from each other through the feature interaction models, thereby significantly improving the effect of medical image segmentation.
[0111] It should be noted that for the functions or steps that the above computer-readable storage medium or computer device can implement, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0112] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0113] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0114] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A medical image segmentation method, characterized in that: The medical image segmentation method comprises: Acquiring medical images; Image segmentation is performed based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image, wherein the medical image segmentation model includes each segmentation model, a feature interaction model corresponding to each segmentation model, and a fusion model.
2. The medical image segmentation method according to claim 1, characterized in that: The segmentation model includes a first segmentation model, a second segmentation model, and a third segmentation model; the feature interaction model includes a first feature interaction model corresponding to the first segmentation model, a second feature interaction model corresponding to the second segmentation model, and a third feature interaction model corresponding to the third segmentation model; and the step of performing image segmentation based on the medical image and the trained medical image segmentation model to obtain a segmentation result of the medical image includes: Inputting the medical image into the first segmentation model to obtain a first feature output by the first segmentation model; Inputting the medical image into the second segmentation model to obtain a second feature output by the second segmentation model; Inputting the medical image into the third segmentation model to obtain a third feature output by the third segmentation model; Image segmentation is performed based on the first feature, the second feature, the third feature, the first feature interaction model, the second feature interaction model, the third feature interaction model and the fusion model to obtain a segmentation result of the medical image.
3. The medical image segmentation method according to claim 2, characterized in that: The step of performing image segmentation based on the first feature, the second feature, the third feature, the first feature interaction model, the second feature interaction model, the third feature interaction model and the fusion model to obtain a segmentation result of the medical image includes: Determine a fourth feature based on the first feature, the second feature, the third feature, and the first feature interaction model; Determine a fifth feature based on the first feature, the second feature, the third feature, and the second feature interaction model; Determine a sixth feature based on the first feature, the second feature, the third feature, and the third feature interaction model; Image segmentation is performed based on the fourth feature, the fifth feature, the sixth feature and the fusion model to obtain a segmentation result of the medical image.
4. The medical image segmentation method according to claim 3, characterized in that: The step of determining the fourth feature based on the first feature, the second feature, the third feature and the first feature interaction model comprises: Based on performing a linear transformation on the second feature, a first Key and a first Value are obtained; Based on performing a linear transformation on the third feature, a first Query is obtained; Input the first Key, the first Value, and the first Query into the first attention model in the first feature interaction model to obtain the seventh feature; Inputting the seventh feature into the convolution layer in the first feature interaction model to obtain an eighth feature output by the convolution layer, and performing a linear transformation on the eighth feature to obtain a second Key and a second Value; Based on performing a linear transformation on the first feature, a second Query is obtained; The second Key, the second Value, and the second Query are input into the second attention model in the first feature interaction model to obtain the fourth feature.
5. The medical image segmentation method according to claim 3, characterized in that: The step of determining the fifth feature based on the first feature, the second feature, the third feature and the second feature interaction model includes: Based on performing a linear transformation on the first feature, a third Key and a third Value are obtained; Based on performing a linear transformation on the third feature, a third Query is obtained; Input the third Key, the third Value, and the third Query into the third attention model in the second feature interaction model to obtain a ninth feature; Inputting the ninth feature into the convolution layer in the second feature interaction model to obtain a tenth feature output by the convolution layer, and performing a linear transformation on the tenth feature to obtain a fourth Key and a fourth Value; Based on performing a linear transformation on the second feature, a fourth Query is obtained; The fourth Key, the fourth Value, and the fourth Query are input into the fourth attention model in the second feature interaction model to obtain the fifth feature.
6. The medical image segmentation method according to claim 3, characterized in that: The step of determining the fifth feature based on the first feature, the second feature, the third feature and the second feature interaction model includes: Based on performing a linear transformation on the first feature, a fifth Key and a fifth Value are obtained; Based on performing a linear transformation on the second feature, a fifth Query is obtained; Input the fifth Key, the fifth Value, and the fifth Query into the fifth attention model in the third feature interaction model to obtain the eleventh feature; Input the eleventh feature into the convolution layer in the third feature interaction model to obtain the twelfth feature output by the convolution layer, and obtain the sixth Key and the sixth Value based on performing a linear transformation on the twelfth feature; Based on performing a linear transformation on the third feature, a sixth Query is obtained; The sixth Key, the sixth Value, and the sixth Query are input into the sixth attention model in the third feature interaction model to obtain the sixth feature.
7. The medical image segmentation method according to claim 3, characterized in that: The step of performing image segmentation based on the fourth feature, the fifth feature, the sixth feature and the fusion model to obtain a segmentation result of the medical image comprises: The fourth feature, the fifth feature, and the sixth feature are spliced together to obtain a target feature; Input the target feature into a 3D global average pooling layer in a fusion model to obtain a first vector, and input the first vector into an excitation layer in the fusion model to obtain a second vector; Inputting the target feature into a 3D global maximum pooling layer in a fusion model to obtain a third vector, and inputting the third vector into an excitation layer in the fusion model to obtain a fourth vector; The second vector and the fourth vector are added to obtain a fifth vector, and the fifth vector is activated by an activation function to obtain a probability matrix; The probability matrix is multiplied by the target feature to obtain the segmentation result of the medical image.
8. A medical image segmentation method and device, characterized in that: The medical image segmentation method and device comprises: an acquisition module, used for acquiring a medical image; An image segmentation module is used to perform image segmentation based on the medical image and a trained medical image segmentation model to obtain a segmentation result of the medical image, wherein the medical image segmentation model includes each segmentation model, a feature interaction model corresponding to each segmentation model, and a fusion model.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the medical image segmentation method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the medical image segmentation method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image segmentation method, image segmentation model construction method, image segmentation model construction device and medium
CN116596846A
Image segmentation method and device, equipment and storage medium
CN118015283A
Medical image segmentation method based on feature interaction
CN118134952A
Dynamic multimodal segmentation selection and fusion
US20230386022A1