A multi-modal based cardiac image segmentation method and system
Patent Information
- Application Number
- CN202410162731.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-05
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2044-02-05
AI Technical Summary
[0004](1)缺乏高质量的心脏影像标注数据;
[0027]本发明提供了一种基于多模态的心脏影像分割方法及系统,从多模态医学图像中提取各组学的特征,并使用对抗性训练方法联合正向分割网络和反向映射网络来训练心脏分割算法模型,最后,迁移到跨媒体的心脏影像,从而完成多模态的图像分割,为后续的诊断及并发症预测提供合理的输入数据。
Smart Images

Figure CN118015009B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cardiac image segmentation technology, and particularly relates to a multimodal cardiac image segmentation method and system. Background Technology
[0002] Automated identification of myocardial infarction complications based on medical images is an urgent clinical need. Currently, the diagnosis of myocardial infarction complications mainly relies on physicians' comprehensive evaluation of various clinical and medical imaging data. Cardiac imaging, due to its complexity, variety, and lack of standardized formats, places higher demands on physicians' image interpretation skills, leading to lengthy diagnostic times, low efficiency, and poor repeatability. In special circumstances, higher requirements for physician-patient protection and more extensive pre-screening procedures further reduce the effectiveness of diagnosing myocardial infarction complications.
[0003] Therefore, there is an urgent clinical need to develop automated medical image recognition tools that improve the diagnostic efficiency of myocardial infarction complications through medical-engineering integration, thereby effectively ensuring emergency treatment processes for major public health emergencies and reducing clinical mortality. However, the technical challenges of automated medical image recognition for myocardial infarction complications are as follows:
[0004] (1) Lack of high-quality cardiac imaging annotation data;
[0005] (2) Existing artificial intelligence methods lack rigorous scientific causal reasoning ability;
[0006] (3) There is a lack of comprehensive, in-depth and efficient cooperation between artificial intelligence researchers and cardiovascular clinical experts.
[0007] Artificial intelligence methods include traditional machine learning algorithms and deep learning algorithms. By combining non-invasive diagnostic techniques with visualized data on myocardial infarction and its complications, disease features are extracted, classified, and autonomously learned to gain the ability to automatically identify the disease.
[0008] Artificial intelligence combined with echocardiography primarily diagnoses myocardial infarction and its complications by assessing cardiac systolic function and identifying ventricular wall motion. Tabassian M et al. used independent principal component analysis (PCA) to extract spatiotemporal features of myocardial segmental strain and constructed a K-nearest neighbor-based machine learning classifier to automatically assess early systolic function in myocardial infarction. The model achieved an average classification accuracy greater than 85%, with overall efficiency exceeding that of manual diagnosis. Kusunose K et al. applied deep learning algorithms using deep convolutional neural networks (DCNN) to automatically identify abnormal ventricular wall motion and segmental distribution, achieving diagnostic accuracy close to or even better than that of cardiovascular specialists.
[0009] The application of cardiac magnetic resonance imaging (MRI) to identify myocardial infarction tissue has increasingly become a research hotspot in artificial intelligence. Chinese scholar Zhang used deep learning algorithms combined with cardiac MRI cine imaging to automatically identify myocardial motion characteristics to determine myocardial infarction. Baessler.B. used artificial intelligence methods to extract cardiac texture features from MRI images to distinguish between myocardial infarction groups and control groups. Multivariate logistic regression showed an area under the curve of 0.92, achieving a relatively ideal recognition effect.
[0010] The key to diagnosing myocardial infarction and its complications using artificial intelligence combined with coronary CT lies in the morphological analysis of the stenotic segment of the coronary artery and the detection and quantitative assessment of atherosclerotic plaques. Doeberitz et al. used a machine learning-based algorithm for CT fractional flow reserve, which significantly improved diagnostic efficacy compared to simply using the percentage of vascular stenosis (areas under the curve were 0.93 and 0.61, respectively). Kang.D. et al. applied deep learning algorithms to automatically classify CCTA obstructive and non-obstructive coronary artery diseases with an accuracy of 94%. In China, Shenzhen Keya Medical has developed "Coronary Fractional Flow Reserve Calculation Software" based on CCTA, utilizing deep learning algorithms and medical image analysis technology. This represents a technological innovation in non-invasive coronary function assessment and is the first AI product in China to receive a Class III medical device certificate.
[0011] Overall, research on the application of artificial intelligence in image recognition of myocardial infarction and its complications has shown initial success. However, limited by inconsistent image data formats, still-imperfect machine learning models, and limited interdisciplinary collaborations, current medical-engineering interdisciplinary achievements are still some distance from true clinical application. Future exploration will focus more on optimizing artificial intelligence algorithms, improving model interpretability, researching analysis techniques suitable for multimedia data, and developing more suitable medical platforms. Summary of the Invention
[0012] To overcome the shortcomings of the prior art, this invention provides a multimodal cardiac image segmentation method and system, which can be used for medical detection images of various modalities such as MRI, ultrasound, and CT to complete multimodal image segmentation, providing reasonable input data for subsequent diagnosis and complication prediction, and further ensuring high detection accuracy and speed.
[0013] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0014] The first aspect of this invention provides a multimodal cardiac image segmentation method.
[0015] A multimodal cardiac image segmentation method includes the following steps:
[0016] Acquiring multimodal medical images;
[0017] The preprocessed image is input into a heart segmentation algorithm model that includes a forward segmentation network and a reverse mapping network. The forward segmentation network is used to predict heart segmentation, and then the reverse mapping network is used to recover the original image from the previous forward segmentation network. The heart segmentation algorithm model is trained based on an adversarial training method to obtain a trained heart segmentation algorithm model.
[0018] The trained cardiac segmentation algorithm model is transferred to cross-media cardiac images to complete multimodal image segmentation.
[0019] A second aspect of the present invention provides a multimodal cardiac image segmentation system.
[0020] A multimodal cardiac image segmentation system, comprising:
[0021] The image acquisition module is configured to acquire multimodal medical images;
[0022] The model building and training module is configured to: input the preprocessed image into a heart segmentation algorithm model containing a forward segmentation network and a reverse mapping network, use the forward segmentation network to predict heart segmentation, then use the reverse mapping network to recover the original image from the previous forward segmentation network, and train the heart segmentation algorithm model based on an adversarial training method to obtain a trained heart segmentation algorithm model.
[0023] The transfer module is configured to transfer the trained cardiac segmentation algorithm model to cross-media cardiac images to complete multimodal image segmentation.
[0024] A third aspect of the present invention provides a computer-readable storage medium having a program stored thereon that, when executed by a processor, implements the steps of the multimodal cardiac image segmentation method as described in the first aspect of the present invention.
[0025] A fourth aspect of the present invention provides an electronic device including a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the multimodal cardiac image segmentation method as described in the first aspect of the present invention.
[0026] The above one or more technical solutions have the following beneficial effects:
[0027] This invention provides a multimodal cardiac image segmentation method and system. It extracts features from various omics from multimodal medical images, and uses an adversarial training method to train a cardiac segmentation algorithm model by combining a forward segmentation network and a reverse mapping network. Finally, it transfers the model to cross-media cardiac images to complete multimodal image segmentation, providing reasonable input data for subsequent diagnosis and complication prediction.
[0028] Compared with existing technologies, the deep and shallow features of the present invention have their own significance: the deeper the network, the larger the receptive field, and the network focuses on global features; shallow networks focus more on local features such as texture; edge features are retrieved through feature concatenation.
[0029] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0030] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0031] Figure 1 This is a flowchart of the method in the first embodiment.
[0032] Figure 2 This is a structural diagram of the heart segmentation algorithm model for the first embodiment.
[0033] Figure 3 This is a flowchart of the data processing for the heart segmentation algorithm model in the first embodiment. Detailed Implementation
[0034] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0035] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0036] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0037] Example 1
[0038] This embodiment discloses a multimodal cardiac image segmentation method.
[0039] like Figure 1 , Figure 2 , Figure 3 As shown, a multimodal cardiac image segmentation method includes the following steps:
[0040] Acquiring multimodal medical images;
[0041] The preprocessed image is input into a heart segmentation algorithm model that includes a forward segmentation network and a reverse mapping network. The forward segmentation network is used to predict heart segmentation, and then the reverse mapping network is used to recover the original image from the previous forward segmentation network. The heart segmentation algorithm model is trained based on an adversarial training method to obtain a trained heart segmentation algorithm model.
[0042] The trained cardiac segmentation algorithm model is transferred to cross-media cardiac images to complete multimodal image segmentation.
[0043] This embodiment can be used for medical imaging modalities such as MRI, ultrasound, and CT, ensuring high detection accuracy and speed, and enabling intelligent, rapid, and accurate detection. The specific technical solution is as follows:
[0044] A multimodal cardiac image segmentation algorithm includes the following steps:
[0045] This study fully utilizes existing cardiac MRI, CT, and ultrasound medical image data to extract features from multimodal medical images. An adversarial training method is then used to train a cardiac segmentation algorithm model by combining a forward segmentation network and a backward mapping network. Finally, the model is transferred to cross-media cardiac images to complete multimodal image segmentation.
[0046] S1, Data integration, building an image database with annotations and labels;
[0047] S2 is a heart segmentation method constrained by a reverse mapping mechanism.
[0048] The first step is data integration.
[0049] The data integration steps include:
[0050] Construct an image database with annotations and labels: Collect cardiac medical images and perform data preprocessing, select images containing lesion features and annotate them, use image enhancement techniques, and construct a dataset that meets the format requirements and matches the network structure.
[0051] The medical image dataset, consisting of over 8,000 multimodal images provided by the medical institution, was divided into training, testing, and validation sets. Lesions were identified according to the medical knowledge provided by the institution. Approximately 30% of the images in the dataset were used to segment the lesion regions using the Labelme tool as the training set, 30% of the images were used as the testing set to evaluate the model's generalization ability, and 40% of the images were used as the validation set to adjust the model's hyperparameters and to conduct a preliminary evaluation of the model's capabilities.
[0052] The second step employs a heart segmentation method constrained by a back-mapping mechanism, introducing the concepts of heart segmentation and its prior distribution, as well as heart images and their prior distribution. A forward network is then constructed to predict heart segmentation, and based on this, a back-mapping network is further built to recover the original image from the previous forward network.
[0053] The following content corresponds to the instruction manual appendix. Figure 1 In the section on building feature extraction models, the generative model content within machine learning includes:
[0054] Let I represent heart segmentation, p(I) represent the prior distribution of segmentation, and x represent the heart image, with q(x) representing the prior distribution of the heart image. The general approach is to construct a feedforward network G1: x→I to predict heart segmentation. Here, we further construct a back-mapping network G2: I→x to recover the original image x from I.
[0055] Generally, q can be estimated using parameterized maximum likelihood estimation. φ (I|x) and p θ (x|I) Learning G1 and G2:
[0056] G1(x; φ) = argmax φ q φ (I|x)
[0057] G2(I; θ) = argmax θ p θ (x|I)
[0058] φ and θ represent the parameters learned from the two distributions.
[0059] To ensure that G1 and G2 complement each other during the learning process, consider the following two functions:
[0060] (1) f:X^→X^,f(x)=G2(G1(x)), which means that G1 first obtains the prediction I from x, and then G2 reconstructs x from the obtained I.
[0061] (2) g:I^→I^,g(I)=G1(G2(I)), which means that G2 first obtains the reconstructed x from I, and then G1 predicts I from the reconstructed x.
[0062] Its mathematical description can be expressed as:
[0063] f(x; θ, φ) = argmax x p θ (x|argmax I q φ (I|x))
[0064] g(x; θ, φ) = argmaxI p φ (I|argmax x p θ (x|I))
[0065] The learning of f and g can be treated as a problem of joint distribution matching. First, consider the joint distribution learned by the feedforward network, then consider the joint distribution q learned by the back-end network. φ (x,I)=q φ (I|x)q(x), and then consider the joint distribution p learned by the reverse network. θ (x,I)=p θ (x|I)p(I).
[0066] We can optimize the network by jointly training these two distributions using Kullback-Leibler (KL) divergence learning. The mathematical description is as follows:
[0067]
[0068]
[0069] Where p θ (x)=∫p θ (x,I)p(I)dI,q φ (I)=∫q φ (I,x)q(x)dx.
[0070] Next, two discriminator networks T are applied. ψ1 (x,I) and T ψ2 (x,I) fits the joint distribution generated by the network.
[0071] The following two optimization objectives can be obtained:
[0072]
[0073]
[0074] Where σ(t)=(1+e -t ) -1 This represents the sigmoid function.
[0075] Meanwhile, given ψ1 and ψ2, we can optimize φ and θ:
[0076] minφ, maxψ1, ψ2Eq φ (I|x)(-T ψ1 (x,I))+logp θ (x|I*)+Ep θ (x|I)(-Tψ2 (x,I))+logq φ (I|x)
[0077] When the min-max reaches Nash equilibrium, and the learned network parameters are {φ*,θ*,ψ1*,ψ2*}, we can obtain q. φ* (x,I)≈p θ* (x, I). This indicates that the knowledge learned by the back-end network is well shared with the forward network.
[0078] like Figure 1 As shown, prior knowledge of the disease is obtained from cardiac electrical / mechanical mechanisms and clinicopathological data; by annotating the acquired cardiac imaging data, multimodal image segmentation data for cardiac segmentation algorithm models used for forward segmentation and backward mapping are obtained.
[0079] Next, a cross-media knowledge graph was constructed, and a feature extraction model was built: a variational autoencoder was designed, a decoder was constructed, and a heart segmentation algorithm model was generated. The model was used to extract features from multimodal image data until the model could make stable and independent disease expressions, thereby completing the automatic recognition task. Multiple models were trained to complete the task.
[0080] Finally, the trained model is applied to clinical research. Clinical data is compared with the industry gold standard by physician diagnosis and the artificial intelligence-assisted diagnosis provided in this embodiment, and then combined with the expert group to improve the overall accuracy and efficiency of diagnosis.
[0081] In the model, a forward network and a backward mapping network were constructed based on a dense network. The parameters learned from MRI and echocardiogram images of these two networks were fine-tuned and then transferred to CT images to achieve multimodal cardiac segmentation.
[0082] In step 2, the network architecture can be summarized as follows:
[0083] The network architecture consists of a symmetrical encoder and decoder:
[0084] Input layer: Accepts input images.
[0085] Encoder: Composed of convolutional and pooling layers, used to progressively reduce the size of the image and extract features.
[0086] Decoder: Composed of convolutional layers and upsampling layers, it is used to progressively restore the size of the image and combine the features extracted by the encoder with the corresponding layers of the decoder.
[0087] Skip connections: combine features from the encoder with the corresponding layer in the decoder to restore the resolution of the image.
[0088] Output layer: Outputs a segmented image with the same size as the input image, where each pixel is labeled as one of the segmentation categories.
[0089] The steps can be summarized as follows:
[0090] The first layer of processing takes a 572×572×1 image as input, performs convolution using 64 3×3 convolution kernels, and obtains 64 570×570×1 feature channels through the ReLU function. Then, it performs convolution using 64 3×3 convolution kernels and obtains 64 568×568×1 feature channels through the ReLU function. This is the result of the first layer of processing.
[0091] In the downsampling process, a 2×2 pooling kernel operation is applied to the processing result of the first layer, downsampling the image to half its original size: 284×284×64. 128 convolutional kernels are then used to further extract features, resulting in a new feature image. The above steps are repeated to downsample the new feature image. Each layer undergoes two convolutions to extract image features. With each downsampling layer, the image size is halved, and the number of convolutional kernels is doubled. The final downsampled result is 28×28×1024, meaning there are a total of 1024 feature layers, with each layer having a feature size of 28×28.
[0092] The upsampling process starts from the bottom right corner, deconvolving the 28×28×1024 feature matrix with 512 2×2 convolutional kernels to expand the matrix to 56×56×512. To reduce data loss, the downsampled images on the left are cropped to the same size and then stitched together to add feature layers before convolution to extract features. Each layer undergoes two convolutions to extract features. Each upsampling layer doubles the image size and halves the number of convolutional kernels. The right side, from bottom to top, involves four upsampling processes. In the final step, two 1×1 convolutional kernels are used to convert the 64 feature channels into two, resulting in a final 388×388×2 matrix. This is a binary classification operation, dividing the image into background and target categories.
[0093] Example 2
[0094] This embodiment discloses a multimodal cardiac image segmentation system.
[0095] A multimodal cardiac image segmentation system, comprising:
[0096] The image acquisition module is configured to acquire multimodal medical images;
[0097] The model building and training module is configured to: input the preprocessed image into a heart segmentation algorithm model containing a forward segmentation network and a reverse mapping network, use the forward segmentation network to predict heart segmentation, then use the reverse mapping network to recover the original image from the previous forward segmentation network, and train the heart segmentation algorithm model based on an adversarial training method to obtain a trained heart segmentation algorithm model.
[0098] The transfer module is configured to transfer the trained cardiac segmentation algorithm model to cross-media cardiac images to complete multimodal image segmentation.
[0099] Example 3
[0100] The purpose of this embodiment is to provide a computer-readable storage medium.
[0101] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the multimodal cardiac image segmentation method as described in Embodiment 1 of this disclosure.
[0102] Example 4
[0103] The purpose of this embodiment is to provide an electronic device.
[0104] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the multimodal cardiac image segmentation method as described in Embodiment 1 of this disclosure.
[0105] The steps and methods involved in the apparatuses of Embodiments 2, 3, and 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.
[0106] Those skilled in the art will understand that the modules or steps of the present invention described above can be implemented using general-purpose computer devices. Optionally, they can be implemented using computer-executable program code, thereby allowing them to be stored in a storage device for execution by a computer device, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. The present invention is not limited to any particular combination of hardware and software.
[0107] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A multimodal cardiac image segmentation method, characterized in that, Includes the following steps: Acquiring multimodal medical images; The preprocessed image is input into a heart segmentation algorithm model that includes a forward segmentation network and a reverse mapping network. The forward segmentation network is used to predict heart segmentation, and then the reverse mapping network is used to recover the original image from the previous forward segmentation network. The heart segmentation algorithm model is trained based on an adversarial training method to obtain a trained heart segmentation algorithm model. The trained cardiac segmentation algorithm model is transferred to cross-media cardiac images to complete multimodal image segmentation. The network architecture of the forward segmentation network and the reverse mapping network includes an input layer, an encoder, a decoder, a skip connection layer, and an output layer: The preprocessed image is input into the heart segmentation algorithm model, and the image size is gradually reduced and features are extracted using the convolutional and pooling layers of the encoder. The image size is gradually restored using the convolutional and upsampling layers of the decoder, and the features extracted by the encoder are combined with the corresponding layers of the decoder through skip connections. The final output is a segmented image with the same size as the input image, and each pixel in the segmented image is labeled as one of the segmentation categories; Let I represent the segmentation of the heart. ( ) represents the prior distribution of the segmentation. Indicates cardiac imaging. ( () represents the prior distribution of cardiac images: Construct a feedforward network G1: → To predict heart segmentation, the inverse mapping network G2: → from Restore the original image ; Through parameterized maximum likelihood estimation and Studying G1 and G2: and This represents the parameters learned from the two distributions; Consider the following two functions: This indicates that G1 first started from The prediction was obtained. Then obtained from G2 China Reconstruction ; This indicates that G2 first started from China was rebuilt Then G1 will rebuild it. China's prediction ; Its mathematical description is as follows: Will and The learning is treated as a problem of joint distribution matching.
2. The multimodal cardiac image segmentation method as described in claim 1, characterized in that: Acquire cardiac medical images, including cardiac MRI, CT, and ultrasound, perform data preprocessing, select and label images containing lesion features, perform image enhancement, and construct a dataset.
3. The multimodal cardiac image segmentation method as described in claim 1, characterized in that, Learning pairs by calculating KL divergence and The two distributions are jointly trained, which can be mathematically described as follows: in .
4. The multimodal cardiac image segmentation method as described in claim 3, characterized in that, Applying two discriminator networks 1( , )and 2( , By fitting the joint distribution generated by the network, we obtain the following two optimization objectives: in, This represents the sigmoid function; In the given 1 and 2. Optimize and : When min-max reaches Nash equilibrium, the learned network parameters are { },get At this point, it indicates that the knowledge learned by the backward network has been well shared with the forward network.
5. A multimodal cardiac image segmentation system, characterized in that: include: The image acquisition module is configured to acquire multimodal medical images; The model building and training module is configured to: input the preprocessed image into a heart segmentation algorithm model containing a forward segmentation network and a reverse mapping network, use the forward segmentation network to predict heart segmentation, then use the reverse mapping network to recover the original image from the previous forward segmentation network, and train the heart segmentation algorithm model based on an adversarial training method to obtain a trained heart segmentation algorithm model. The transfer module is configured to transfer the trained cardiac segmentation algorithm model to cross-media cardiac images to complete multimodal image segmentation. The network architecture of the forward segmentation network and the reverse mapping network includes an input layer, an encoder, a decoder, a skip connection layer, and an output layer: The preprocessed image is input into the heart segmentation algorithm model, and the image size is gradually reduced and features are extracted using the convolutional and pooling layers of the encoder. The image size is gradually restored using the convolutional and upsampling layers of the decoder, and the features extracted by the encoder are combined with the corresponding layers of the decoder through skip connections. The final output is a segmented image with the same size as the input image, and each pixel in the segmented image is labeled as one of the segmentation categories; Let I represent the segmentation of the heart. ( ) represents the prior distribution of the segmentation. Indicates cardiac imaging. ( () represents the prior distribution of cardiac images: Construct a feedforward network G1: → To predict heart segmentation, the inverse mapping network G2: → from Restore the original image ; Through parameterized maximum likelihood estimation and Studying G1 and G2: and This represents the parameters learned from the two distributions; Consider the following two functions: This indicates that G1 first started from The prediction was obtained. Then obtained from G2 China Reconstruction ; This indicates that G2 first started from China was rebuilt Then G1 will rebuild it. China's prediction ; Its mathematical description is as follows: Will and The learning is treated as a problem of joint distribution matching.
6. A computer-readable storage medium having a program stored thereon, characterized in that, When executed by a processor, the program implements the steps in the multimodal cardiac image segmentation method as described in any one of claims 1-4.
7. An electronic device, comprising a memory, a processor, and a program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the multimodal cardiac image segmentation method as described in any one of claims 1-4.
Citation Information
Patent Citations
Myocardial infarction complication identification method based on heart multi-mode image
CN120071059A