Medical image segmentation method and device, electronic equipment and storage medium

By performing weighted fusion and complementary mask processing on multimodal MRI images, the accuracy problem of the multimodal MRI image segmentation model is solved, and a more accurate image segmentation effect is achieved.

CN120689606APending Publication Date: 2025-09-23SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510573397.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Existing medical image segmentation models have difficulty in effectively processing multimodal MRI images, especially the correlation information between T1-weighted images, T2-weighted images, enhanced T1-weighted images and FLAIR images, resulting in inaccurate image segmentation results.

Method used

By performing image weighted fusion on MRI images of multiple modalities, determining the dominant modality and secondary modality, and performing complementary mask processing, a pre-built initial image feature extractor is used to extract features. Combined with image reconstruction and model training, a target image feature extractor is formed to improve image segmentation accuracy.

Benefits of technology

It enhances the understanding of feature associations between multimodal medical images and improves the accuracy and stability of image segmentation, especially in medical image segmentation scenarios where some content is missing, and can also achieve more accurate segmentation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689606A_ABST
    Figure CN120689606A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a medical image segmentation method and device, electronic equipment and a storage medium, and relates to the technical field of image processing. The method comprises the following steps: acquiring original sample medical images of different modalities; determining a dominant mode and a secondary mode; performing image weighted fusion on the original sample medical image and the dominant weight corresponding to the dominant mode and the original sample medical image and the secondary weight corresponding to the secondary mode to obtain a mixed mode image; performing complementary mask processing on the mixed modal image to obtain a mixed image mask; performing feature extraction on each mixed image mask through a pre-constructed initial image feature extractor to obtain initial image features; and performing model training on the initial image feature extractor according to the initial image features and the original sample medical image corresponding to the dominant mode to obtain a target image feature extractor so as to obtain an image segmentation model. According to the embodiment of the invention, the image segmentation precision of the multi-modal medical image can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a medical image segmentation method and device, electronic equipment, and storage medium. Background Art

[0002] Medical image segmentation is a technique for dividing medical images into regions with distinct semantic features (such as anatomical or pathological characteristics). For example, in tumor diagnosis and treatment, to clearly define the boundaries of tumors in human organs, such as the boundaries of gliomas in the brain, brain magnetic resonance images (MRI images) can be segmented to identify the tumor region in the MRI image.

[0003] Currently, image segmentation models based on deep neural networks are widely used in medical image segmentation scenarios. However, current image segmentation models are often limited to processing single-modality medical images (such as CT images and MRI images of the same type) and are difficult to use for processing multi-modality MRI images (including T1-weighted images, T2-weighted images, enhanced T1-weighted images, FLAIR images, etc.). For example, MRI images of different modalities (such as T1-weighted images and T2-weighted images) are usually input independently into the image segmentation model. The image segmentation model cannot understand the relationship between MRI images of different modalities, which makes the image segmentation results less accurate.

[0004] Therefore, how to improve the accuracy of image segmentation of multimodal medical images has become a technical problem that needs to be solved urgently. Summary of the Invention

[0005] The main purpose of the embodiments of the present application is to propose a medical image segmentation method and device, electronic device, and storage medium, aiming to understand the correlation information between multimodal medical images, thereby improving the accuracy of image segmentation of multimodal medical images.

[0006] To achieve the above objectives, a first aspect of an embodiment of the present application provides a medical image segmentation method, the method comprising:

[0007] Acquire at least two original sample medical images of the sample organ; wherein any two of the original sample medical images correspond to different modalities;

[0008] determining each of at least two of the modes as a dominant mode in turn, and determining the modes other than the dominant mode as secondary modes;

[0009] Performing image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight;

[0010] Performing complementary mask processing on the mixed-modality images corresponding to at least two of the dominant modalities to obtain a mixed-image mask corresponding to each of the dominant modalities; wherein mask regions of the mixed-image masks of at least two of the dominant modalities are complementary;

[0011] Performing feature extraction on the mixed image mask corresponding to each of the dominant modes using a pre-built initial image feature extractor to obtain initial image features corresponding to each of the dominant modes;

[0012] Performing image reconstruction on the initial image features corresponding to each of the dominant modalities to obtain a reconstructed medical image;

[0013] Performing model training on the initial image feature extractor according to the reconstructed medical image corresponding to each of the dominant modalities and the original sample medical image corresponding to the dominant modality to obtain a target image feature extractor;

[0014] A pre-trained image segmentor is connected after the target image feature extractor to obtain an image segmentation model for performing image segmentation on medical images.

[0015] To achieve the above-mentioned purpose, a second aspect of an embodiment of the present application provides a medical image segmentation device, comprising:

[0016] An image acquisition module, configured to acquire at least two original sample medical images of a sample organ; wherein any two of the original sample medical images correspond to different modalities;

[0017] a mode determination module, configured to sequentially determine each of the at least two modes as a dominant mode, and determine the modes other than the dominant mode as secondary modes;

[0018] an image weighted fusion module, configured to perform image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight;

[0019] a complementary mask module, configured to perform complementary masking on the mixed-modality images corresponding to at least two of the dominant modalities to obtain a mixed-image mask corresponding to each of the dominant modalities; wherein the mask regions of the mixed-image masks of at least two of the dominant modalities are complementary;

[0020] a feature extraction module, configured to extract features from the mixed image mask corresponding to each of the dominant modes using a pre-built initial image feature extractor, to obtain initial image features corresponding to each of the dominant modes;

[0021] an image reconstruction module, configured to reconstruct the initial image features corresponding to each of the dominant modalities to obtain a reconstructed medical image;

[0022] a model training module, configured to perform model training on the initial image feature extractor based on the reconstructed medical image corresponding to each of the dominant modalities and the original sample medical image corresponding to the dominant modality, to obtain a target image feature extractor;

[0023] A model building module is used to connect a pre-trained image segmentor after the target image feature extractor to obtain an image segmentation model for image segmentation of medical images.

[0024] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the method described in the first aspect when executing the computer program.

[0025] To achieve the above-mentioned purpose, the fourth aspect of the embodiments of the present application proposes a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the method described in the first aspect.

[0026] The medical image segmentation method and device, electronic device, and storage medium proposed in this application obtain original sample medical images corresponding to different modalities, determine each of the multiple modalities as the dominant modality in turn, and then perform image weighted fusion on the original sample medical images corresponding to each modality, and the weight corresponding to the dominant modality (i.e., the dominant weight) is greater than the weight corresponding to the secondary modality (i.e., the secondary weight), thereby fully fusing the multimodal medical images (i.e., the original sample medical images) and ensuring the dominant position of each dominant modality in the corresponding mixed modality image, so that the initial image feature extractor can more comprehensively understand the characteristics of the medical images corresponding to each modality during the model training process. Then, the mixed modality image is subjected to complementary masking processing, which ensures that the distribution of the masked area in medical images of different modalities is different, so that the initial image feature extractor can be trained based on the cross-modal information between multiple medical images, so that the target image feature extractor obtained after training can understand the feature associations between multimodal medical images to more accurately extract image features, such as the anatomical structure of the target organ, the style characteristics of each modality, etc., so that the image segmentation model can more accurately understand the semantic features contained in the multimodal medical image, thereby improving the accuracy of image segmentation of multimodal medical images. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a flowchart of the medical image segmentation method provided in an embodiment of the present application;

[0028] Figure 2 This is a schematic diagram of a specific implementation of the medical image segmentation method provided in an embodiment of the present application;

[0029] Figure 3 yes Figure 1 Flowchart of step 102 in FIG.

[0030] Figure 4 yes Figure 3 Flowchart of step 205 in FIG.

[0031] Figure 5 yes Figure 1 Flowchart of step 104 in FIG.

[0032] Figure 6 yes Figure 5 Flowchart of step 403 in FIG.

[0033] Figure 7 is a flowchart of a medical image segmentation method provided by another embodiment of the present application;

[0034] Figure 8 is a flowchart of a medical image segmentation method provided by another embodiment of the present application;

[0035] Figure 9 is a structural diagram of a medical image segmentation device provided in an embodiment of the present application;

[0036] Figure 10 This is a schematic diagram of the hardware structure of the electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0037] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0038] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.

[0039] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0040] First, let’s analyze some of the terms used in this application:

[0041] Artificial Intelligence (AI) is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. A branch of computer science, AI attempts to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thinking. It can also refer to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. Basic AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning. This application can acquire and process relevant data based on AI technologies.

[0042] Nuclear Magnetic Resonance Imaging (NMRI): Also known as Magnetic Resonance Imaging (MRI), MRI is a medical imaging technology based on the principle of nuclear magnetic resonance (NMR). It uses the energy absorption and release characteristics of hydrogen nuclei (protons) in the human body under the influence of strong magnetic fields and radio frequency pulses to generate images by detecting the electromagnetic wave signals released by them.

[0043] Weighted image: Also known as weighted image, refers to an image in medical imaging that emphasizes or suppresses specific tissue signals by adjusting imaging parameters. In MRI, weighted images are formed based on the signal differences between different tissues under specific physical parameters (such as T1 relaxation time and T2 relaxation time). By setting pulse sequence parameters (such as repetition time TR and echo time TE), the signals of different tissues are assigned different weights, and an image that emphasizes the characteristic parameters of a certain tissue can be obtained. This image is called a weighted image. Weighted images in MRI include T1-weighted images, T2-weighted images, etc.

[0044] Repetition time (TR): refers to the time interval between successive pulse sequences applied to the same slice. Specifically, TR is the time interval between the application of one radiofrequency (RE) pulse and the application of the next RF pulse.

[0045] Echo time (TE) is the time interval between the emission of the RF pulse and the acquisition of the echo signal. Specifically, the echo time is the time interval from the start of the RF pulse to the receipt of the echo signal (also known as the magnetic resonance signal).

[0046] T1-weighted image (T1WI): An MRI image that emphasizes differences in longitudinal relaxation time (T1 relaxation time) between tissues. T1-weighted images are characterized by short TR and TE.

[0047] T2-weighted image (T2WI): An MRI image that emphasizes differences in transverse relaxation time (T2 relaxation time) between tissues. T2-weighted images are characterized by long TR and TE.

[0048] Enhanced-T1 Weighted Image (E-T1WI): refers to the use of specific technical means to enhance the contrast or signal intensity of T1-weighted images, thereby showing the MRI images of specific tissues or lesions more clearly. The main difference between enhanced T1-weighted images and conventional T1-weighted images is whether contrast agents are used. T1-weighted images highlight the longitudinal relaxation differences of tissues by adjusting TR and TE parameters, while enhanced T1-weighted images further enhance this contrast effect by injecting contrast agents. Specifically, in the process of generating enhanced T1-weighted images, the number of hydrogen protons in the tissue is changed by injecting contrast agents, thereby affecting the T1 relaxation time and enhancing the contrast of the image.

[0049] FLAIR images are MRI images obtained using a fluid-attenuated inversion recovery (FLAIR) sequence. In FLAIR images, the signal of free water (such as cerebrospinal fluid) is suppressed while the high signal of bound water (such as water molecules in diseased tissue) is retained.

[0050] The medical image segmentation method and device, electronic device, and storage medium provided in the embodiments of the present application are specifically illustrated through the following embodiments. First, the medical image segmentation method in the embodiments of the present application is described.

[0051] The medical image segmentation method provided in the embodiments of the present application can be applied to a terminal, can be applied to a server, and can also be software running on a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server can be configured as an independent physical server, or as a server cluster or distributed system consisting of multiple physical servers, or as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms; the software can be an application that implements the medical image segmentation method, etc., but is not limited to the above forms.

[0052] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0053] Figure 1 This is an optional flowchart of the medical image segmentation method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps 101 to 108.

[0054] Step 101: Acquire at least two original sample medical images of a sample organ; wherein the modalities corresponding to any two original sample medical images are different;

[0055] Step 102 , determining each of the at least two modes as a dominant mode in turn, and determining modes other than the dominant mode as secondary modes;

[0056] Step 103: performing image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight;

[0057] Step 104: performing complementary mask processing on the mixed-modality images corresponding to at least two dominant modalities to obtain a mixed-image mask corresponding to each dominant modality; wherein the mask regions of the mixed-image masks of at least two dominant modalities are complementary;

[0058] Step 105: extract features from the mixed image mask corresponding to each dominant mode using a pre-built initial image feature extractor to obtain initial image features corresponding to each dominant mode;

[0059] Step 106, performing image reconstruction on the initial image features corresponding to each dominant modality to obtain a reconstructed medical image;

[0060] Step 107: performing model training on the initial image feature extractor based on the reconstructed medical image corresponding to each dominant modality and the original sample medical image corresponding to the dominant modality to obtain a target image feature extractor;

[0061] Step 108 : Connect a pre-trained image segmenter after the target image feature extractor to obtain an image segmentation model for performing image segmentation on medical images.

[0062] The beneficial effects of the embodiments of the present application include but are not limited to: by obtaining original sample medical images corresponding to different modalities, each modality in multiple modalities is determined as the dominant modality in turn, and then the original sample medical images corresponding to each modality are subjected to image weighted fusion, and the weight corresponding to the dominant modality (i.e., the dominant weight) is greater than the weight corresponding to the secondary modality (i.e., the secondary weight), thereby fully fusing the multimodal medical images (i.e., the original sample medical images) and ensuring the dominant position of each dominant modality in the corresponding mixed modality image, so that the initial image feature extractor can more comprehensively understand the characteristics of the medical images corresponding to each modality during the model training process. Then, the mixed modality image is subjected to complementary masking processing, which ensures that the distribution of the masked area in medical images of different modalities is different, so that the initial image feature extractor can be trained based on the cross-modal information between multiple medical images, so that the target image feature extractor obtained after training can understand the feature associations between multimodal medical images to more accurately extract image features, such as the anatomical structure of the target organ, the style characteristics of each modality, etc., so that the image segmentation model can more accurately understand the semantic features contained in the multimodal medical image, thereby improving the accuracy of image segmentation of multimodal medical images.

[0063] In step 101 of some embodiments, the sample organ refers to an organ corresponding to the original sample medical image, such as the brain or spine. For example, if each of the original sample medical images is a brain MRI image obtained by performing magnetic resonance imaging (MRI) on the brain, then the sample organ corresponding to the original sample medical image is the brain. In another embodiment, the sample organ may also include other types of organs, without limitation.

[0064] It should be noted that original sample medical images refer to medical images (also known as medical images) acquired from a sample organ, such as MRI images obtained by performing magnetic resonance imaging (MRI) of the sample organ. Any two original sample medical images correspond to different modalities. Therefore, multiple original sample medical images can form a multimodal medical image set to train the initial image feature extractor's feature extraction capabilities for multimodal medical images.

[0065] In some embodiments, when the original sample medical images are MRI images, any two original sample medical images correspond to different modalities, which means that any two original sample medical images correspond to different nuclear magnetic resonance sequences. It should be noted that a nuclear magnetic resonance sequence refers to an imaging method that obtains different tissue contrast and information by combining different parameters such as radio frequency pulses, gradient magnetic fields, and signal acquisition time in a magnetic resonance imaging examination. For example, Figure 2 As shown, the multiple input images (i.e., original sample medical images) may include T1-weighted images (hereinafter referred to as weighted images), T2-weighted images, enhanced T1-weighted images, and FLAIR images. The modality corresponding to the T1-weighted image is the T1 modality, i.e., the T1-weighted imaging (T1WI) sequence. The modality corresponding to the T2-weighted image is the T2 modality, i.e., the T2-weighted imaging (T2WI) sequence. The modality corresponding to the enhanced T1-weighted image is the T1 ce modality, i.e., the T1-weighted enhanced imaging (T1 ce) sequence. The modality corresponding to the FLAIR image is the FLAIR modality, i.e., the FLAIR sequence.

[0066] In some embodiments, before step 102, each original sample medical image may be subjected to image enhancement processing to improve image quality for subsequent image processing. For example, each original sample medical image may be subjected to random spatial transformation, such as any one or more of random image cropping, image scaling, and image rotation.

[0067] In step 102 of some embodiments, each of the modalities corresponding to the original sample medical images is sequentially determined as the dominant modality. For example, assume there are two original sample medical images, including sample medical image A and sample medical image B, where sample medical image A corresponds to modality A and sample medical image B corresponds to modality B. Modality A can be initially determined as the dominant modality, and modality B as the secondary modality, thereby obtaining a mixed-modality image in which modality A is dominant during the subsequent weighted image fusion process. Subsequently, modality B can be determined as the dominant modality, and modality A as the secondary modality, thereby obtaining a mixed-modality image in which modality B is dominant.

[0068] In step 103 of some embodiments, the dominant weight refers to the weight used to multiply the original sample medical images corresponding to the dominant modality. The secondary weight refers to the weight used to multiply the original sample medical images corresponding to the secondary modality. It should be noted that the number of mixed modality images is the same as the number of modalities, and the number of mixed modality images is also the same as the number of original sample medical images. For example, Figure 2 As shown in , if there are 4 original sample medical images, there are 4 modalities and 4 mixed modality images. Specifically, each original sample medical image corresponds to a modality, and each modality is dominant in a mixed modality image.

[0069] In step 104 of some embodiments, it should be noted that the mask areas of the mixed image masks of at least two dominant modes are complementary, which means that the visible areas (i.e., mask areas) of any two mixed image masks are different and the visible areas of each mixed image mask are complementary. Figure 2 In , T1 mask represents the mixed image mask of the mixed modality image corresponding to T1 modality, T2 mask represents the mixed image mask of the mixed modality image corresponding to T2 modality, T1 ce mask represents the mixed image mask of the mixed modality image corresponding to T1 ce modality, and FLAIR mask represents the mixed image mask of the mixed modality image corresponding to FLAIR modality. Figure 2 In the four masks shown above, the black squares represent invisible areas, and the other areas (such as the four gray areas in the T1 mask) represent visible areas. The visible areas of each channel (i.e., the mixed image mask) are disjoint and complementary.

[0070] It should be noted that since each mixed-modality image has a corresponding dominant modality, complementary masking between images (also known as complementary masking) can be roughly viewed as complementary masking between modalities. Complementary masking can destroy the mixed image, thereby helping to train the initial image feature encoder to understand the anatomical features, semantic features, and other feature information contained in medical images of different modalities. For example, a damaged image (i.e., a mixed image mask) is obtained through complementary masking operations, and then the original medical image is restored from the damaged image. This training method is more robust, is less affected by noise in the image, and can preserve the anatomical details contained in the medical image.

[0071] In step 105 of some embodiments, the initial image feature extractor is a model for extracting features from an image. Specifically, the initial image feature extractor can be based on an encoder-decoder (see Figure 2 The initial image feature extractor may also be other types of models, but is not limited thereto.

[0072] In step 106 of some embodiments, image reconstruction can be performed by an image reconstructor (also called a reconstruction head). The image reconstructor refers to a deep learning network model used for image reconstruction. The number of reconstructed medical images is consistent with the number of original sample medical images. For example, Figure 2 As shown, there are 4 reconstruction outputs, that is, 4 reconstructed medical images, each of which corresponds to a dominant modality.

[0073] In step 107 of some embodiments, the target image feature extractor refers to the trained initial image feature extractor. Specifically, a loss calculation can be performed based on the reconstructed medical image of the dominant modality and the original sample medical image corresponding to the dominant modality to obtain an image reconstruction loss value. The initial image feature extractor is then trained based on the image reconstruction loss value to obtain the target image feature extractor.

[0074] It should be noted that while the image reconstruction loss value is used to train the initial image feature extractor, the image reconstruction loss value can also be used to train the reconstruction head.

[0075] In some embodiments, after obtaining the initial image features corresponding to each dominant modality, a predicted mixing coefficient sequence can be determined for the initial image features and the original sample medical image through a weight predictor (also called a prediction head). The predicted mixing coefficient sequence refers to the predicted weight combination used when the original sample medical image is fused into the mixed modality image corresponding to the initial image features. Then, a loss calculation is performed based on the difference between the predicted mixing coefficient sequence and the actual target mixing coefficient sequence (including the dominant weights and the secondary weights) to obtain a mixing coefficient prediction loss value; the initial image feature extractor is model trained based on the mixing coefficient prediction loss value to obtain a target image feature extractor.

[0076] It should be noted that while the mixing coefficient prediction loss value is used to train the initial image feature extractor, the mixing coefficient prediction loss value can also be used to train the prediction head.

[0077] It should be noted that in Figure 2 In the figure, the pink squares represent the target mixing coefficient sequence when the T1 mode is the dominant mode. The green squares represent the target mixing coefficient sequence when the T2 mode is the dominant mode. The yellow squares represent the target mixing coefficient sequence when the T1ce mode is the dominant mode. The blue-purple squares represent the target mixing coefficient sequence when the FLAIR mode is the dominant mode.

[0078] In some embodiments, for example, the initial image feature extractor may be trained using only image reconstruction or only mixing coefficient prediction. For another example, the initial image feature extractor may be trained using both of the above methods simultaneously, which is not limited in the present embodiment.

[0079] In step 108 of some embodiments, the image segmenter (also called reconstruction head) may be a model based on image semantic segmentation, such as a U-Net model.

[0080] In some embodiments, after obtaining the image segmentation model, a sample modality medical image and its corresponding sample label image can also be obtained; based on the sample modality medical image and the sample label image, the image segmentation model is fine-tuned to perform image segmentation on the medical image corresponding to the target organ through the fine-tuned image segmentation model.

[0081] In some embodiments, the medical image segmentation method of the embodiment of the present application can be applied to brain tumor image segmentation. It should be noted that brain tumors (especially gliomas) are one of the neurological diseases with the highest mortality rate. Accurate image segmentation of brain tumor images has high clinical value. For example, accurate tumor segmentation is a key step in formulating treatment plans (such as surgical resection range, radiotherapy target area delineation) and evaluating prognosis. By performing pixel-level or voxel-level segmentation on medical images (such as MRI images), the boundary of the tumor, the infiltration range of the tumor, and the anatomical relationship between the tumor and the surrounding tissue can be clarified, thereby improving the diagnostic accuracy and improving the patient's quality of life. Since MRI images of different modalities can reflect the pathological characteristics of the tumor (such as edema, necrosis, degree of enhancement, etc.) from different dimensions, MRI images of multiple modalities (including T1-weighted images, T2-weighted images, enhanced T1-weighted images, FLAIR images, etc.) are widely used in clinical practice.

[0082] However, current image segmentation models struggle to extract feature information from multiple modalities of medical images (referred to as multimodal medical images) and accurately segment them. For example, in the field of medical image segmentation, current image segmentation models based on deep neural networks (such as U-Net and Transformer architectures) are typically trained using supervised learning. However, supervised learning training relies on a large amount of high-quality labeled sample data. Medical images must be annotated pixel by pixel by professional radiologists, a manual process that consumes significant time and labor costs. Furthermore, the high subjectivity of manual annotation results in poor quality of the resulting labeled sample data. Furthermore, the extensive medical expertise required to annotate medical images, the complexity of 3D medical images further increasing the cost of annotation, and the scarcity of medical images related to rare diseases (such as certain brain tumors) all contribute to the limited availability of labeled sample data. Consequently, image segmentation models trained using supervised learning have poor generalization capabilities and are often applied to images of a single modality, making it difficult to accurately segment multimodal medical images. For another example, a self-supervised learning training method can be used to train image segmentation models. However, image segmentation models trained through self-supervised learning are usually applied to single-modality medical images (such as CT images) and are difficult to use for processing multi-modality MRI images. For example, different medical images are usually directly regarded as independent channel input models for model training. Therefore, image segmentation models, especially image feature extractors in image segmentation models, find it difficult to understand cross-modal correlation information, such as the associated anatomical features of T1-weighted images and FLAIR images of the same brain tumor.

[0083] Taking the above-mentioned deficiencies into consideration, the embodiment of the present application enhances the data diversity of the samples by performing image weighted fusion on multimodal medical images, so that the initial image feature extractor can learn the features of each modality more fully from the samples. Moreover, through the operation of complementary masks, the initial image feature extractor learns the complementary information of different modalities and establishes cross-modal anatomical correspondences. This can improve the accuracy, efficiency and stability of image segmentation. In addition, since the image segmentation model has the ability to understand the global features of multiple different modalities during the training process, it can also obtain more accurate segmentation results in medical image segmentation scenarios where there is a modality missing (such as image segmentation of MRI images with missing content).

[0084] See also Figure 3 In some embodiments, step 103 may include but is not limited to steps 201 to 207:

[0085] Step 201: performing region division on the original sample medical image corresponding to the dominant modality to obtain at least two dominant image regions, and performing region division on the original sample medical image corresponding to the secondary modality to obtain at least two candidate secondary image regions; wherein positions of the at least two candidate secondary image regions correspond one-to-one to positions of the at least two dominant image regions;

[0086] Step 202 , for each dominant image region, selecting a region at the same position as the dominant image region from at least two candidate secondary image regions to obtain a target secondary image region;

[0087] Step 203: Determine the number of at least two modes as the modality number, and divide 1 by the modality number to obtain an average weight threshold;

[0088] Step 204: Obtain a target mixing coefficient sequence; wherein the target mixing coefficient sequence includes mixing weights equal to the number of modes, and the sum of the mixing weights equal to the number of modes is 1;

[0089] Step 205: Select a mixing weight greater than an average weight threshold from the at least two mixing weights to obtain a dominant weight, and determine the mixing weights other than the dominant weight as secondary weights;

[0090] Step 206 , performing weighted fusion based on the product of the dominant weight and the dominant image area, and the product of each secondary weight and each target secondary image area, to obtain a mixed image area;

[0091] Step 207 : performing image region combination on at least two mixed image regions to obtain a mixed modality image.

[0092] The advantage of this embodiment is that by dividing the original sample medical image corresponding to the dominant modality into dominant image regions and the original sample medical image corresponding to the secondary modality into candidate secondary image regions, a weighted fusion is then performed based on the product of the dominant weight and the dominant image region, and the product of each secondary weight and each target secondary image region, to obtain a mixed image region, thereby improving the mixing adequacy of each modality in the mixed image (mixed modality image). Specifically, the number of at least two modalities is determined as the modality number, and 1 is divided by the modality number to obtain an average weight threshold; a target mixing coefficient sequence including the number of modalities, the mixing weight, is obtained from the at least two mixing weights, and the mixing weight greater than the average weight threshold is selected from the at least two mixing weights to obtain the dominant weight, and the mixing weights other than the dominant weight are determined as secondary weights. This allows for more complete mixing of the modalities and ensures the dominant position of the dominant modality in the mixed modality image.

[0093] In step 201 of some embodiments, the number of modalities refers to the number of modalities corresponding to the at least two element sample medical images, and the number of modalities is consistent with the number of original sample medical images.

[0094] In step 202 of some embodiments, for example, Figure 2 As shown, if the dominant image area is located at the 3rd row and the 1st column, the target secondary image area is also located at the 3rd row and the 1st column.

[0095] In step 203 of some embodiments, the average weight threshold is a ratio of 1 to the number of modalities. The average weight threshold is used to represent the average weight corresponding to each modality. For example, if the number of modalities is 4, the average weight threshold is 1 / 4 = 0.25.

[0096] In step 204 of some embodiments, the number of mixing weights is equivalent to the number of modalities, and each mixing weight corresponds to one modality.

[0097] In step 205 of some embodiments, the dominant weight is greater than the average weight threshold, and the dominant weight corresponds to the dominant modality, thereby ensuring the dominant position of the dominant modality in the mixed modality image. In some embodiments, the dominant weight is greater than 0.

[0098] In step 206 of some embodiments, the dominant weight is a weight used for weighted fusion of the dominant image region, and the secondary weight is a weight used for weighted fusion of the target secondary image region.

[0099] In step 207 of some embodiments, the mixed image regions may be arranged into a mixed modality image according to the positions (including row coordinates and column coordinates) of the mixed image regions.

[0100] See also Figure 4 In some embodiments, step 204 may include but is not limited to steps 301 to 304:

[0101] Step 301: Generate an integer sequence for consecutive integers from 1 to the modal number to obtain a target integer sequence; wherein the target integer sequence includes at least two target integers;

[0102] Step 302: Based on each target integer, weight sampling is performed on at least two preset candidate weights to obtain a target integer number of target weights; wherein each target weight is greater than 0, and the sum of the target integer number of target weights is 1;

[0103] Step 303: If the target integer is less than the number of modes, the difference between the number of modes and the target integer is determined as the number of zero-padding, and the combination of the target integer number of target weights and the number of zero-padding zeros is determined as a candidate mixing coefficient sequence;

[0104] Step 304: Randomly select at least two candidate mixing coefficient sequences to obtain a target mixing coefficient sequence.

[0105] The advantage of this embodiment is that, based on each consecutive integer from 1 to the number of modes, at least two preset candidate weights are weight sampled to obtain a target integer number of target weights to generate a candidate mixing coefficient sequence, and then a target mixing coefficient sequence is randomly selected from the at least two candidate mixing coefficient sequences. This improves the randomness of obtaining the target mixing coefficient sequence, thereby increasing the diversity and robustness of the training samples used for model training of the initial image feature extractor, thereby improving the reliability of image segmentation.

[0106] In some embodiments, taking the number of modes as 4 as an example, a mixing coefficient sequence containing 1, 2, 3, and 4 weights greater than 0 (i.e., the target weights mentioned above) can be sampled from the preset candidate numerical sequence {0.05, 0.1, 0.15, ..., 1}, and the sum of all target weights in the mixing coefficient sequence is equal to 1. The candidate numerical sequence is a sequence from 0 to 1 with an interval of 0.05, and the values ​​in the sequence are greater than 0. For example, a mixing coefficient sequence L1 = {1} containing 1 target weight, a mixing coefficient sequence L2 = {0.55, 0.45} containing 2 weights greater than 0, a mixing coefficient sequence L3 = {0.4, 0.35, 0.25} containing 3 weights greater than 0, and a mixing coefficient sequence L4 = {0.4, 0.3, 0.2, 0.1} containing 4 weights greater than 0 can be sampled from the candidate numerical sequence. In addition, for mixing coefficient sequences with fewer weights than the number of modes, the missing weights are padded with 0 to obtain a candidate mixing coefficient sequence. For example, the candidate mixing coefficient sequence L1' after padded with 0 = {1, 0, 0, 0}, the candidate mixing coefficient sequence L2' = {0.55, 0.45, 0, 0}, and the candidate mixing coefficient sequence L3' = {0.4, 0.35, 0.25, 0}. The candidate mixing coefficient sequence L4' does not need to be padded with 0, so the candidate mixing coefficient sequence L4' is the same as the mixing coefficient sequence L4. Then, a sequence is randomly selected from the above four candidate mixing coefficient sequences, which is the target mixing coefficient sequence mentioned above.

[0107] In step 301 of some embodiments, the target integer sequence includes all integers from 1 to the number of modes, i.e., target integers. For example, if the number of modes is 4, the target integer sequence includes four target integers, specifically 1, 2, 3, and 4. The target integers correspond to the number of weights greater than 0 in the target mixing coefficient sequence.

[0108] It should be noted that, through computer exhaustive enumeration, there are 1495 weight combinations (i.e., candidate mixing coefficient sequences) for the four modalities that meet the constraints that the sum of the four weights is 1 and the dominant weight is greater than 0. Among them, one weight combination contains one weight greater than 0, meaning that the mixed-modality image obtained by weighted image fusion based on this weight combination can retain the characteristic content of one modality (i.e., the dominant modality); 54 weight combinations contain two weights greater than 0, meaning that the mixed-modality image obtained by weighted image fusion based on this weight combination can retain the characteristic content of two modalities (i.e., the dominant modality and one secondary modality); 501 weight combinations contain three weights greater than 0; and 939 weight combinations contain four weights greater than 0. Therefore, if random sampling of the mixing coefficient sequence is performed directly under the above constraints, more than 90% of the target mixing coefficient sequences obtained by sampling will include 3 or more weights greater than 0. In other words, the mixed-modal image will contain feature content of 3 or more modalities, which will make it difficult to ensure the dominant position of the dominant modality in the mixed-modal image. Therefore, the embodiment of the present application first generates one candidate mixing coefficient sequence each including 1, 2, 3, and 4 weights greater than 0, and then randomly selects the 4 candidate mixing coefficient sequences to obtain the target mixing coefficient sequence. This can better ensure the dominant position of the dominant modality in the mixed-modal image, so that in subsequent model training, the initial image feature extractor can more accurately learn the features of the dominant modality.

[0109] In step 302 of some embodiments, for normalization, the sum of the target integer target weights is 1. Similarly to the above example, when the target integer is 2, the sum of the two target weights in the candidate mixing coefficient sequence L2′={0.55, 0.45, 0, 0} is 0.55+0.45=1.

[0110] In step 303 of some embodiments, the zero-padding amount is the difference between the modal amount and the target integer.

[0111] In step 304 of some embodiments, the target mixing coefficient sequence is one of at least two candidate mixing coefficient sequences.

[0112] See also Figure 5 In some embodiments, step 104 may include, but is not limited to, steps 401 to 404:

[0113] Step 401: determining the number of at least two modes as the number of modes, and determining the product of the number of modes and a preset target multiple as the number of segmented regions;

[0114] Step 402 , performing average region division on each mixed modality image according to the number of segmentation regions to obtain a number of candidate unit regions for the segmentation regions;

[0115] Step 403: Select a target multiple of candidate unit regions from the candidate unit regions corresponding to each dominant modality to obtain a target multiple of target unit regions; wherein each mixed modality image corresponds to a different target unit region;

[0116] Step 404 : Determine a combination of target unit areas that are multiples of the target as a mixed image mask.

[0117] The advantage of this embodiment is that by setting the number of segmented regions to a multiple of the number of modalities, each mixed-modality image is evenly divided into regions, and each mixed-image mask is then ensured to include a target multiple of the target unit regions. This ensures that the regions obtained after segmenting the mixed-modality image corresponding to each dominant modality are evenly distributed to the mixed-image mask corresponding to that dominant modality, so that the mask regions of each mixed-image mask have the same area. As a result, the amount of feature information retained by each mixed-image mask is relatively similar, allowing the initial image feature extractor to more comprehensively learn the feature associations between the various modalities.

[0118] In step 401 of some embodiments, the number of segmented regions is a multiple of the number of modalities.

[0119] In step 402 of some embodiments, for example, Figure 2 As shown, each mixed modality image can be divided into 16 candidate unit areas in an average manner according to a 4×4 division method.

[0120] In step 403 of some embodiments, the candidate unit regions may be randomly selected under different constraints for the target unit regions corresponding to each mixed-modality image.

[0121] In step 404 of some embodiments, for example, Figure 2 As shown, when the target multiple is 4, each mixed image mask includes 4 target unit areas, and the 4 target unit areas together constitute the visible area of ​​the mixed image mask.

[0122] See also Figure 6 ,In some embodiments, the target multiple is 4, and each candidate unit area has row coordinates and column coordinates;

[0123] Step 403 may include but is not limited to steps 501 to 503:

[0124] Step 501 , selecting two candidate unit regions with different row coordinates from at least four candidate unit regions corresponding to each dominant mode to obtain a first unit region and a second unit region;

[0125] Step 502: Delete the first unit area and the second unit area from at least four candidate unit areas;

[0126] Step 503 : Select two candidate unit areas with different column coordinates from at least two candidate unit areas to obtain a third unit area and a fourth unit area; wherein the target unit area includes the first unit area, the second unit area, the third unit area, and the fourth unit area.

[0127] The advantage of this embodiment is that by selecting unit areas in different rows, such as the first and second unit areas, and selecting unit areas in different columns, such as the third and fourth unit areas, the distribution of each target unit area in the mixed image mask can be made more dispersed, that is, the mask area is made more dispersed and covers the global information of the image. This makes it possible to infer the missing content of the image based on the global information of the medical image to reconstruct the image, thereby reducing the local dependence of the target feature extractor on the image, and further improving the image segmentation stability of the image segmentation model composed of the target feature extractor.

[0128] In step 501 of some embodiments, the row coordinates of the first unit area and the second unit area are different.

[0129] In step 502 of some embodiments, the first unit area and the second unit area are deleted from at least four candidate unit areas, so that the first unit area or the second unit area can be repeatedly selected from the candidate unit areas later.

[0130] In step 503 of some embodiments, the column coordinates of the third unit area and the fourth unit area are different.

[0131] See also Figure 7 In some embodiments, before step 103, the medical image segmentation method may further include but is not limited to steps 601 to 603:

[0132] Step 601, performing a random linear geometric transformation on each original sample medical image to obtain an enhanced medical image;

[0133] Step 602: performing similarity calculation based on the enhanced medical image and the original sample medical image to obtain image similarity;

[0134] Step 603: If the image similarity is greater than or equal to a preset similarity threshold, the enhanced medical image is determined as the original sample medical image.

[0135] The advantage of this embodiment is that random linear geometric transformations, such as random cropping, scaling, and rotation, are performed on each original sample medical image to simulate the natural differences in medical images in clinical scenarios (such as changes in patient position and differences in equipment parameters), thereby enhancing the diversity of model training data and enabling the initial image feature extractor to more fully learn image features that are not affected by spatial position and scale, thereby improving the generalization ability of the image segmentation model composed of the target image feature extractor. In addition, considering that the process of random linear geometric transformation may cause key information in the original medical image (such as tumor-related information) to be defective or even lost, the similarity between the transformed image (i.e., the enhanced medical image) and the original image (i.e., the original sample medical image) is first calculated. Only when the similarity is greater than a preset similarity threshold is the enhanced medical image used for subsequent processing. In this way, the stability of the model in image segmentation of medical images can be improved, and it is suitable for scenarios where complex medical images (such as brain MRI images) are segmented.

[0136] In step 601 of some embodiments, the random linear geometric transformation may include one or more of random cropping, random scaling, and random rotation. For example, random scaling can be used to simulate medical images of organs at different scanning resolutions. Random rotation can also be used to simulate angular deviations during medical image acquisition. In another embodiment, the random linear geometric transformation may also include other operations, not limited to these.

[0137] In step 602 of some embodiments, the image similarity may be calculated using a cosine similarity algorithm or a histogram comparison algorithm, which is not limited in the embodiments of the present application.

[0138] In step 603 of some embodiments, the similarity threshold may be 70%. The similarity threshold may also be set or adjusted to other values ​​as required.

[0139] See also Figure 8 In some embodiments, after step 105, the medical image segmentation method may further include but is not limited to steps 701 to 702:

[0140] Step 701, performing mask image prediction on the initial image features to obtain a predicted mixed image;

[0141] Step 702: performing weight prediction based on the predicted mixed image and at least two original sample medical images to obtain at least two prediction weights;

[0142] Step 703 : Perform model training on the initial image feature extractor based on at least two prediction weights, the dominant weight, and the secondary weight to obtain a target image feature extractor.

[0143] The advantage of this embodiment is that the initial image feature extractor is trained by predicting weights, thereby improving the target image feature extractor's ability to understand the feature associations between different modalities, and further improving the accuracy of image segmentation performed by the image segmentation model composed of the target image feature extractor.

[0144] In step 701 of some embodiments, a masked image prediction may be performed on the initial image features using a Masked Image Modeling (MIM) algorithm to obtain a predicted mixed image.

[0145] In step 702 of some embodiments, the at least two prediction weights include a dominant prediction weight and a secondary prediction weight. Weight prediction can be performed using a deep neural network model or other methods.

[0146] In step 703 of some embodiments, a mixed coefficient prediction loss value can be calculated based on the similarity between the dominant prediction weights and the similarity between the secondary prediction weights to train the initial image feature extractor.

[0147] See also Figure 9 The present invention also provides a medical image segmentation device that can implement the above-mentioned medical image segmentation method. The device includes:

[0148] An image acquisition module 801 is configured to acquire at least two original sample medical images of a sample organ; wherein the modalities corresponding to any two original sample medical images are different;

[0149] a mode determination module 802 for determining each of the at least two modes as a dominant mode in sequence, and determining modes other than the dominant mode as secondary modes;

[0150] The image weighted fusion module 803 is configured to perform image weighted fusion on the original sample medical images corresponding to the dominant modality and the dominant weight, and on the original sample medical images corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight;

[0151] A complementary masking module 804 is configured to perform complementary masking on the mixed-modality images corresponding to at least two dominant modalities to obtain a mixed-image mask corresponding to each dominant modality; wherein the mask regions of the mixed-image masks of at least two dominant modalities are complementary;

[0152] A feature extraction module 805 is configured to extract features from the mixed image mask corresponding to each dominant modality using a pre-built initial image feature extractor to obtain initial image features corresponding to each dominant modality;

[0153] An image reconstruction module 806 is configured to reconstruct the initial image features corresponding to each dominant modality to obtain a reconstructed medical image;

[0154] A model training module 807 is configured to perform model training on the initial image feature extractor based on the reconstructed medical image corresponding to each dominant modality and the original sample medical image corresponding to the dominant modality to obtain a target image feature extractor;

[0155] The model building module 808 is used to connect a pre-trained image segmentor after the target image feature extractor to obtain an image segmentation model for performing image segmentation on medical images.

[0156] In one embodiment, the medical image segmentation device also includes an image enhancement module, which is used to: perform a random linear geometric transformation on each original sample medical image to obtain an enhanced medical image; calculate similarity between the enhanced medical image and the original sample medical image to obtain image similarity; if the image similarity is greater than or equal to a preset similarity threshold, determine the enhanced medical image as the original sample medical image.

[0157] In one embodiment, the medical image segmentation device also includes a weight prediction module, which is used to: perform mask image prediction on the initial image features to obtain a predicted mixed image; perform weight prediction based on the predicted mixed image and at least two original sample medical images to obtain at least two predicted weights; and perform model training on the initial image feature extractor based on the at least two predicted weights, the dominant weight and the secondary weight to obtain a target image feature extractor.

[0158] The specific implementation of the medical image segmentation device is basically the same as the specific embodiment of the above-mentioned medical image segmentation method, and will not be repeated here.

[0159] The present application also provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned medical image segmentation method when executing the computer program. The electronic device may include any smart terminal such as a tablet computer or an in-vehicle computer.

[0160] See also Figure 10 , Figure 10 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes:

[0161] The processor 901 can be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application.

[0162] The memory 902 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the medical image segmentation method of the embodiments of this application.

[0163] Input / output interface 903, used to implement information input and output;

[0164] Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.);

[0165] Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 );

[0166] The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0167] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the above-mentioned medical image segmentation method is implemented.

[0168] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0169] It should be noted that the non-Company's software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

[0170] The embodiments described in this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0171] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0172] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0173] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0174] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0175] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0176] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0177] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0178] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0179] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0180] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A medical image segmentation method, characterized in that: The method comprises: Acquire at least two original sample medical images of the sample organ; wherein any two of the original sample medical images correspond to different modalities; determining each of at least two of the modes as a dominant mode in turn, and determining the modes other than the dominant mode as secondary modes; Performing image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight; Performing complementary mask processing on the mixed-modality images corresponding to at least two of the dominant modalities to obtain a mixed-image mask corresponding to each of the dominant modalities; wherein mask regions of the mixed-image masks of at least two of the dominant modalities are complementary; Performing feature extraction on the mixed image mask corresponding to each of the dominant modes using a pre-built initial image feature extractor to obtain initial image features corresponding to each of the dominant modes; Performing image reconstruction on the initial image features corresponding to each of the dominant modalities to obtain a reconstructed medical image; Performing model training on the initial image feature extractor according to the reconstructed medical image corresponding to each of the dominant modalities and the original sample medical image corresponding to the dominant modality to obtain a target image feature extractor; A pre-trained image segmentor is connected after the target image feature extractor to obtain an image segmentation model for performing image segmentation on medical images.

2. The method according to claim 1, characterized in that The performing image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality, includes: performing region division on the original sample medical image corresponding to the dominant modality to obtain at least two dominant image regions, and performing region division on the original sample medical image corresponding to the secondary modality to obtain at least two candidate secondary image regions; wherein positions of the at least two candidate secondary image regions correspond one-to-one to positions of the at least two dominant image regions; For each of the dominant image regions, selecting a region at the same position as the dominant image region from at least two of the candidate secondary image regions to obtain a target secondary image region; Determine the number of at least two of the modes as the number of modes, and divide 1 by the number of modes to obtain an average weight threshold; Obtaining a target mixing coefficient sequence; wherein the target mixing coefficient sequence includes mixing weights of the number of modes, and the sum of the mixing weights of the number of modes is 1; Selecting the mixing weight greater than the average weight threshold from at least two of the mixing weights to obtain the dominant weight, and determining the mixing weights other than the dominant weight as the secondary weights; Perform weighted fusion according to the product of the dominant weight and the dominant image area, and the product of each of the secondary weights and each of the target secondary image areas, to obtain a mixed image area; Image region combination is performed on at least two of the mixed image regions to obtain the mixed modality image.

3. The method according to claim 2, characterized in that The obtaining of the target mixing coefficient sequence includes: Generating an integer sequence for consecutive integers from 1 to the modal number to obtain a target integer sequence; wherein the target integer sequence includes at least two target integers; Based on each of the target integers, weight sampling is performed on at least two preset candidate weights to obtain the target integer target weights; wherein each of the target weights is greater than 0, and the sum of the target integer target weights is 1; If the target integer is less than the number of modes, a difference between the number of modes and the target integer is determined as the number of zero-padding, and a combination of the target integer number of target weights and the number of zeros in the zero-padding is determined as a candidate mixing coefficient sequence; At least two candidate mixing coefficient sequences are randomly selected to obtain the target mixing coefficient sequence.

4. The method according to any one of claims 1 to 3, characterized in that The performing complementary mask processing on the mixed modality images corresponding to at least two of the dominant modalities to obtain a mixed image mask corresponding to each of the dominant modalities includes: Determining the number of at least two of the modes as the mode number, and determining the product of the mode number and a preset target multiple as the number of segmented regions; Performing average region division on each of the mixed modality images according to the number of segmented regions to obtain a number of candidate unit regions of the segmented regions; Selecting the target multiple of the candidate unit areas from the candidate unit areas corresponding to each of the dominant modalities to obtain the target multiple of the target unit areas; wherein the target unit areas corresponding to each of the mixed modal images are different; A combination of the target unit areas of the target multiple is determined as the mixed image mask.

5. The method according to claim 4, characterized in that The target multiple is 4, and each candidate unit area has row coordinates and column coordinates; The step of selecting the target multiple of the candidate unit areas from the candidate unit areas corresponding to each of the dominant modes to obtain the target multiple of the target unit areas includes: Selecting two candidate unit regions with different row coordinates from the at least four candidate unit regions corresponding to each dominant mode to obtain a first unit region and a second unit region; deleting the first unit area and the second unit area from at least four candidate unit areas; From at least two of the candidate unit areas, two candidate unit areas with different column coordinates are selected to obtain a third unit area and a fourth unit area; wherein the target unit area includes the first unit area, the second unit area, the third unit area and the fourth unit area.

6. The method according to any one of claims 1 to 3, characterized in that Before performing image weighted fusion on the original sample medical images corresponding to the dominant modality and the dominant weights, and on the original sample medical images corresponding to the secondary modalities and the secondary weights to obtain mixed modality images corresponding to each dominant modality, the method further includes: Performing a random linear geometric transformation on each of the original sample medical images to obtain an enhanced medical image; performing similarity calculation based on the enhanced medical image and the original sample medical image to obtain image similarity; If the image similarity is greater than or equal to a preset similarity threshold, the enhanced medical image is determined as the original sample medical image.

7. The method according to any one of claims 1 to 3, characterized in that After extracting features from the mixed image mask corresponding to each of the dominant modes using a pre-built initial image feature extractor to obtain initial image features corresponding to each of the dominant modes, the method further includes: Performing mask image prediction on the initial image features to obtain a predicted mixed image; Performing weight prediction based on the predicted mixed image and at least two original sample medical images to obtain at least two prediction weights; The initial image feature extractor is trained according to at least two of the prediction weights, the dominant weight, and the secondary weight to obtain the target image feature extractor.

8. A medical image segmentation device, characterized in that: The device comprises: An image acquisition module, configured to acquire at least two original sample medical images of a sample organ; wherein any two of the original sample medical images correspond to different modalities; a mode determination module, configured to sequentially determine each of the at least two modes as a dominant mode, and determine the modes other than the dominant mode as secondary modes; an image weighted fusion module, configured to perform image weighted fusion on the original sample medical image corresponding to the dominant modality and the dominant weight, and on the original sample medical image corresponding to the secondary modality and the secondary weight, to obtain a mixed modality image corresponding to each dominant modality; wherein the dominant weight is greater than the secondary weight; a complementary mask module, configured to perform complementary masking on the mixed-modality images corresponding to at least two of the dominant modalities to obtain a mixed-image mask corresponding to each of the dominant modalities; wherein the mask regions of the mixed-image masks of at least two of the dominant modalities are complementary; a feature extraction module, configured to extract features from the mixed image mask corresponding to each of the dominant modes using a pre-built initial image feature extractor, to obtain initial image features corresponding to each of the dominant modes; an image reconstruction module, configured to reconstruct the initial image features corresponding to each of the dominant modalities to obtain a reconstructed medical image; a model training module, configured to perform model training on the initial image feature extractor based on the reconstructed medical image corresponding to each of the dominant modalities and the original sample medical image corresponding to the dominant modality, to obtain a target image feature extractor; A model building module is used to connect a pre-trained image segmentor after the target image feature extractor to obtain an image segmentation model for image segmentation of medical images.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the medical image segmentation method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the medical image segmentation method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Multi-modal medical image segmentation method and device, electronic equipment and medium

    CN121354228A