Cross-modal heart image segmentation method and system based on space-time causal

The potential variables are extracted through the spatiotemporal causal diffusion network model, which solves the problem of insufficient spatiotemporal causal connections in traditional cross-modal heart image segmentation, and realizes the precise segmentation of heart images and the accurate provision of structural information, supporting the diagnosis and treatment of heart diseases.

CN120259653AInactive Publication Date: 2025-07-04ZHENGZHOU UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510315623.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When the traditional cross-modal heart image segmentation method processes heart images of different modalities, it is difficult to effectively capture the causal relationship of the heart in the space-time dimension, resulting in insufficient segmentation accuracy and stability and cannot meet clinical needs.

Method used

The cross-modal heart image segmentation method based on spatiotemporal causality is adopted, and the latent variables are extracted through the spatiotemporal causal diffusion network model using the 3D spatial detangling block, temporal detangling block and modal encoder. The cardiac reconstruction image is generated by combining causal intervention and diffusion modules to achieve accurate segmentation of the heart image.

Benefits of technology

It can accurately segment the heart structure, provide clear cardiac structure information, solve the problem of space-time confusion in cross-modal heart image segmentation, and assist doctors in the diagnosis and treatment of heart disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259653A_ABST
    Figure CN120259653A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-modal heart image segmentation method and system based on space-time causal. The method comprises the following steps: acquiring a clinical heart image; the clinical heart image comprises a magnetic resonance heart image, a computed tomography heart image and an ultrasonic heart image; identifying the image type of the clinical heart image, and obtaining annotation information corresponding to the image type; the image type comprises a time sequence image, a 3D image and a 3D time sequence image; inputting the clinical heart image and the annotation information into the space-time causal diffusion network model to obtain a heart image segmentation result; the space-time causal diffusion network model is used for performing image segmentation on the clinical heart image according to a set mode and a corresponding image type to obtain a heart image segmentation result; the heart image segmentation result comprises a heart segmentation image and annotation data. By adopting the method, the dynamic change characteristics of the heart in different time and space can be captured according to different types of clinical heart images, and a heart image segmentation result is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image region segmentation, and particularly relates to a cross-modal cardiac image segmentation method and system based on spatio-temporal causality. Background Art

[0002] With the development of medical imaging technology, various modal imaging technologies have emerged in the field of cardiac imaging, such as magnetic resonance imaging (MR), computed tomography (CT), and ultrasound imaging (US), providing rich information for the diagnosis and treatment of heart diseases.

[0003] In traditional technologies, the field of cross-modal cardiac image segmentation mainly includes model-based methods. The model-based cross-modal cardiac image segmentation method segments cardiac images by constructing a specific mathematical model.

[0004] However, for traditional cross-modal cardiac image segmentation methods, on the one hand, they have high requirements for the selection of unique features of the segmented images. Manually select key feature points to accurately represent the positional relationship and morphological features of different anatomical structures of the heart. Due to the differences in the morphology, size, and position of the hearts of different patients, the staff selecting the feature points need to understand cardiac anatomy, pathology, and be familiar with professional knowledge such as image features and anatomical relationships. On the other hand, most cross-modal cardiac image segmentation methods are based on invariant representations, and this kind of representation mostly relies on correlation rather than causality, resulting in poor generalization ability of the model between different modalities, being difficult to adapt to the application in clinical scenarios, and also unable to effectively solve the causal relationship of cardiac images in the spatio-temporal and spatial dimensions, making the segmentation accuracy and stability difficult to meet clinical requirements. Summary of the Invention

[0005] Based on this, it is necessary to provide a cross-modal cardiac image segmentation method and system based on spatio-temporal causality for the above technical problems, which can segment cardiac imaging and provide doctors with clear and accurate cardiac structure information related to spatio-temporal association.

[0006] In a first aspect, the present application provides a cross-modal cardiac image segmentation method based on spatio-temporal causality, including:

[0007] Obtain clinical cardiac images; the clinical cardiac images include magnetic resonance cardiac images, computed tomography cardiac images, and ultrasound cardiac images;

[0008] Identify the image type of the clinical cardiac images and obtain the annotation information corresponding to the image type; the image types include time series images, 3D images, and 3D time series images;

[0009] Input the clinical heart images and annotation information into the spatio-temporal causal diffusion network model to obtain the heart image segmentation results; the spatio-temporal causal diffusion network model is used to perform image segmentation on the clinical heart images according to the corresponding image type in a set pattern to obtain the heart image segmentation results; the heart image segmentation results include heart segmentation images and annotation data.

[0010] In one embodiment, performing image segmentation on the clinical heart images according to the corresponding image type in a set pattern to obtain the heart image segmentation results includes:

[0011] Using causal intervention to extract latent variables from the clinical heart images according to the preset pattern of the corresponding image type; the latent variables include anatomical elements and modality elements;

[0012] Generating a heart reconstruction image based on the latent variables;

[0013] Performing convolutional processing on the heart reconstruction image and the anatomical elements to obtain the heart image segmentation results.

[0014] In one embodiment, extracting latent variables from the clinical heart images according to the preset pattern of the corresponding image type includes:

[0015] When the image type is a 3D image, using a 3D spatial unwrapping block to process the clinical heart image to obtain latent variables;

[0016] When the image type is a time series image, using a temporal unwrapping block to process the clinical heart image to obtain latent variables;

[0017] When the image type is a 3D time series image, using the spatial unwrapping block and the temporal unwrapping block together to process the clinical heart image to obtain latent variables.

[0018] In one embodiment, using a 3D spatial unwrapping block to process the clinical heart image to obtain latent variables includes:

[0019] Using the anatomical encoder in the 3D spatial unwrapping block to preliminarily extract 3D anatomical features from the clinical heart image to obtain a 3D anatomical feature map;

[0020] Using the attention mechanism to convert the 3D anatomical feature map into a probability map matching the size of the clinical heart image to obtain an intermediate result containing 3D anatomical features;

[0021] Using the modality encoder in the 3D spatial unwrapping block to perform modality extraction on the clinical heart image to obtain an intermediate result containing CT image modality information;

[0022] Integrating the intermediate result containing 3D anatomical features and the intermediate result containing CT image modality information to obtain latent variables.

[0023] In one embodiment, a latent variable is obtained by processing a clinical cardiac image using a temporal unwrapping block, including:

[0024] Performing parallel processing on a clinical cardiac image through multiple convolutional encoding paths using an anatomical encoder in the temporal unwrapping block to obtain a probability map;

[0025] Performing convolutional processing on the probability map using a convolutional block to convert it into a cached probability map, generating an intermediate result of anatomical elements containing temporal features;

[0026] Performing modality extraction on a clinical cardiac image using a modality encoder in the temporal unwrapping block to obtain an intermediate result containing time series modality information;

[0027] Integrating the intermediate result of anatomical elements containing temporal features and the intermediate result containing time series modality information to obtain a latent variable.

[0028] In one embodiment, a latent variable is obtained by jointly processing a clinical cardiac image using a spatial unwrapping block and a temporal unwrapping block, including:

[0029] Performing preliminary extraction of 3D anatomical features on a clinical cardiac image using an anatomical encoder in a 3D spatial unwrapping block to obtain a 3D anatomical feature map;

[0030] Using an attention mechanism to convert the 3D anatomical feature map into a probability map matching the size of the clinical cardiac image, obtaining an intermediate result containing 3D anatomical features;

[0031] Performing modality extraction on a clinical cardiac image using a modality encoder in a 3D spatial unwrapping block to obtain an intermediate result containing MR image modality information;

[0032] Performing parallel processing on a clinical cardiac image through multiple convolutional encoding paths using an anatomical encoder in the temporal unwrapping block to obtain a probability map;

[0033] Performing convolutional processing on the probability map using a convolutional block to convert it into a cached probability map, generating an intermediate result of anatomical elements containing temporal features;

[0034] Performing modality extraction on a clinical cardiac image using a modality encoder in the temporal unwrapping block to obtain an intermediate result containing time series modality information;

[0035] Fusing the intermediate result containing 3D anatomical features and the intermediate result of anatomical elements containing temporal features to obtain anatomical elements;

[0036] Fusing the intermediate result containing MR image modality information and the intermediate result containing time series modality information to obtain modality elements;

[0037] The anatomical elements and the modality elements constitute a latent variable.

[0038] In one embodiment, generating a cardiac reconstruction image based on latent variables includes:

[0039] Generating an initial image using the latent variables;

[0040] Adding Gaussian noise to the initial image to obtain a noisy image;

[0041] Reshaping the noisy image using Gaussian transfer to obtain a cardiac reconstruction image.

[0042] In a second aspect, the present application also provides a cross-modal cardiac image segmentation system based on spatio-temporal causality, the system includes:

[0043] An image acquisition module for acquiring clinical cardiac images;

[0044] An image recognition module for identifying the image type of the clinical cardiac image and obtaining annotation information corresponding to the image type;

[0045] A model algorithm module for obtaining the clinical cardiac image and the annotation information, and performing image segmentation on the clinical cardiac image according to a set mode according to different image types to obtain a cardiac image segmentation result.

[0046] In one embodiment, the model algorithm module includes a dynamic causality module, a diffusion module, and a segmentation module;

[0047] The dynamic causality module is used to extract latent variables from the clinical cardiac image according to a preset mode corresponding to the image type by using causal intervention; the causal intervention is that the dynamic causality module uses the objective function of the dynamic causality module to reflect the spatio-temporal causal relationship characteristics of the cardiac segmentation image and the annotation data, and extracts latent variables to determine the spatio-temporal causal relationship characteristics of the cardiac segmentation image and the annotation data;

[0048] The diffusion module is used to generate a cardiac reconstruction image according to the latent variables;

[0049] The segmentation module is used to perform convolution processing on the cardiac reconstruction image to obtain a cardiac image segmentation result.

[0050] In one embodiment, when the image type is a 3D image, the following objective function of the dynamic causality module is used to reflect the spatio-temporal causal relationship characteristics of the cardiac segmentation image and the annotation data:

[0051] argminE[Y S |do(X s =x s ),I 0:S-1

[0052]

[0053] ​Among them, argmin is the intervention strategy that minimizes the expected loss, and E[Y S |do(X s =x s ),I 0:S-1 is the expected condition; do(X s =x s ) is the causal intervention, and I 0:S-1 is the expected value of all interventions and the corresponding label Y S for the known previous spatial dimensions; Y S is the labeled data;

[0054] When the image type is a time-series image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship characteristics between the cardiac segmentation image and the labeled data:

[0055] argminE[Y t |do(X t =x t ),I 0:t-1

[0056]

[0057] Among them, argmin is the intervention strategy that minimizes the expected loss, and E[Y t |do(X t =x t ),I 0:t-1 is the expected condition; do(X t =x t ), is the causal intervention, and I 0:t-1 is the expected value of all interventions and the corresponding label Y t for the known previous time dimensions; Y t is the labeled data;

[0058] When the image type is a 3D time-series image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship characteristics between the cardiac segmentation image and the labeled data:

[0059] argminE[Y t,s |do(X t.s =x t,s ),I 0:t-1,s-1

[0060]

[0061] Among them, argmin is the intervention strategy that minimizes the expected loss, and E[Y t,s |do(X t.s =x t,s 0,I 0:t-1,s-1 is the expected condition; do(X​​t.s = x t,s 0 is the causal intervention, I 0:t-1,s-1 is all the interventions I in the known previous time dimension 0:t-1 and all the interventions I in the previous space dimension 0:S-1 and the jointly corresponding label Y t,s the expected value of; Y t,s is the labeled data.

[0062] The above cross-modal cardiac image segmentation method and system based on spatio-temporal causality can analyze the time and space correlation of different types of clinical cardiac images, capture the dynamic change characteristics of the heart in different spatio-temporal domains, so as to accurately segment the cardiac structure, solve the spatio-temporal confusion problem in cross-modal cardiac image segmentation, obtain the segmented image and labeled data, provide clear and accurate cardiac structure information for doctors, directly observe the morphology, size, positional relationship of each part of the structure of the heart through the segmented image, and obtain quantitative cardiac parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 is a schematic flow chart of the cross-modal cardiac image segmentation method based on spatio-temporal causality of the present invention;

[0065] Figure 2 is a sub-step flow chart of step S103;

[0066] Figure 3 is a sub-step flow chart of step S202;

[0067] Figure 4 is a composition structure diagram of the cross-modal cardiac image segmentation system based on spatio-temporal causality of the present invention;

[0068] Figure 5 is a composition structure diagram of the model algorithm module in the system. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0069] In order to make the objectives, technical solutions and advantages of the present application clearer, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0070] In one embodiment, asFigure 1 As shown, a cross-modal cardiac image segmentation method based on spatio-temporal causality is provided. In this embodiment, this method is exemplified by being applied to a terminal. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0071] S101. Obtain clinical cardiac images, where the clinical cardiac images include magnetic resonance cardiac images, computed tomography cardiac images, and echocardiographic images.

[0072] Schematically, clinical cardiac images can be directly obtained from data sources such as hospital imaging databases and imaging devices. The cardiac images corresponding to existing cardiac imaging modalities include magnetic resonance cardiac images, computed tomography cardiac images, and echocardiographic images. Each type of image contains specific information about the heart and can reflect the physiological structure of the heart from different perspectives. Different modalities of cardiac images have their own advantages. Specifically, magnetic resonance cardiac images have high resolution for soft tissues and can clearly show structures such as the myocardium. Computed tomography cardiac images can provide high-resolution images of the cardiac anatomical structure, which helps to observe the morphology and blood vessels of the heart. Echocardiographic images can dynamically display the motion of the heart in real time. These clinical cardiac images provide the original data for the subsequent analysis and segmentation of clinical cardiac images.

[0073] S102. Identify the image type of the clinical cardiac image and obtain the annotation information corresponding to the image type. The image types include time series images, 3D images, and 3D time series images.

[0074] According to the characteristics of different imaging modalities, determine the type of the obtained clinical cardiac image to prepare for subsequent extraction of different features, that is, determine whether the clinical cardiac image is a time series image, a 3D image, or a 3D time series image. Schematically, generally when the clinical cardiac image is a magnetic resonance cardiac image, the corresponding image type is a 3D time series image, which has cardiac information in both time and space dimensions. When the clinical cardiac image is a computed tomography cardiac image, the corresponding image type is a 3D image, presenting the three-dimensional spatial structure of the heart. When the clinical cardiac image is an echocardiographic image, the corresponding image type is a time series image, reflecting the dynamic changes of the heart over a period of time. Different image types require different processing methods and parameters in the follow-up. Identifying the image type helps the subsequent spatio-temporal causality diffusion network model to select an appropriate processing flow.

[0075] S103, inputting the clinical cardiac image and the annotation information into the spatiotemporal causal diffusion network model to obtain a cardiac image segmentation result. The spatiotemporal causal diffusion network model is used to perform image segmentation on the clinical cardiac image according to the set mode and the corresponding image type to obtain a cardiac image segmentation result. The cardiac image segmentation result includes a cardiac segmentation image and annotation data.

[0076] The acquired clinical cardiac images and their corresponding annotation information are input into a pre-trained spatiotemporal causal diffusion network model. Multiple modules within the model process the clinical cardiac images. The processing steps are determined by the image type of the clinical cardiac images. Further, the temporal features, spatial features or spatiotemporal features are extracted in detail according to the image type. The cardiac segmentation image is an image that accurately divides the various parts of the heart structure from the original image. The annotation data is a detailed annotation of the various parts of the structure in the cardiac image segmentation result, and is also the cardiac structure information data on the cardiac segmentation image.

[0077] This method integrates a variety of clinical cardiac images such as magnetic resonance imaging (MR), computed tomography (CT) and ultrasound (US). Each modality image can provide cardiac information from different angles. Different types of images such as time series images, 3D images and 3D time series images are identified, and the corresponding annotation information is obtained. For different image types, the spatiotemporal causal diffusion network model uses a specific setting mode for processing. For time series US images, the model can capture the dynamic change characteristics of the heart in the time dimension and accurately segment the shape of the heart at different times; for 3D CT images, the model focuses on analyzing the three-dimensional spatial structure of the heart and accurately outlines the boundaries of the heart's chambers, myocardium and other structures; for 3D time series MR images, the model can consider both the spatial structure of the heart and its changes over time to achieve more accurate segmentation. The spatiotemporal causal diffusion network model analyzes images based on spatiotemporal causal relationships. During the processing process, the model can dig out the causal factors that really affect the segmentation of cardiac structures in the image through causal intervention and learning from previous intervention information, avoiding incorrect segmentation caused by spatiotemporal confusion. By utilizing the powerful feature extraction and image segmentation capabilities of the spatiotemporal causal diffusion network model, clinical cardiac images can be accurately segmented, and complex cardiac images can be converted into intuitive segmentation results, providing doctors with accurate cardiac structure information to assist them in diagnosing heart diseases and formulating treatment plans.

[0078] In one embodiment, if Figure 2 As shown, according to the set mode, the clinical cardiac image is segmented according to the corresponding image type to obtain the cardiac image segmentation result, including:

[0079] S201. Use causal intervention to extract latent variables from clinical cardiac images according to the preset mode corresponding to the image type; the latent variables include anatomical elements and modal elements.

[0080] Taking the clinical cardiac image as input, taking the MR image as an example, because it has 3D and temporal characteristics, it will be processed by both the 3D spatial disentangling block and the temporal disentangling block at the same time; the CT image is mainly processed by the 3D spatial disentangling block; the US image focuses on the temporal disentangling block. Thus, latent variables are extracted. The latent variables are features that aim to accurately reflect the causal relationship between the clinical cardiac image and the annotation information.

[0081] S202. Generate a cardiac reconstruction image according to the latent variables.

[0082] The latent variables contain various anatomical result features and specific imaging modalities in the clinical cardiac image, including the morphological and positional information of the myocardium, heart, etc. The latent variables go through three different diffusion processes. Each process uses Unet (U-shaped Network) as the backbone network, including a forward process and a reverse process, so as to generate a cardiac reconstruction image.

[0083] S203. Perform convolution processing on the cardiac reconstruction image and the anatomical elements to obtain the cardiac image segmentation result.

[0084] According to the correspondence relationship between the parameters and features obtained by training and the annotation information, process the reconstructed image. Further, for the cardiac reconstruction image of the 3D image, the cardiac reconstruction image and the anatomical elements are processed by two 3D convolution modules. For the cardiac reconstruction image of the time series image, it is processed by a 2D convolution module. Each convolution module includes a convolutional layer, a batch normalization layer, and PReLU (Parametric Rectified Linear Unit, activation function), with the number of channels being 63. Finally, the 3D convolution block outputs the predicted annotation data to complete the segmentation of the clinical cardiac image and output the cardiac image segmentation result. Specifically, the predicted annotation data is obtained through the following formula:

[0085] y i = h(ni)

[0086] where h represents the depth convolution operation of the 3D convolution block, and n i is an element in the anatomical element N. By performing convolution processing on the anatomical element N, the features corresponding to the structures of each part of the heart can be captured, and then the labels corresponding to different regions in the cardiac image can be predicted, such as the positions of the endocardium and epicardium.

[0087] The cross-modal cardiac image segmentation method based on spatiotemporal causality provided in the embodiment of the present application can deeply explore the causal features closely related to the cardiac structure and function in clinical cardiac images by adopting causal intervention to extract latent variables according to preset patterns of different image types. Therefore, the cardiac image segmentation results can clearly display the different structures of the heart, such as the endocardium, epicardium, etc., to assist doctors in diagnosing and analyzing heart diseases, which is a key operation that directly serves clinical applications.

[0088] In one embodiment, latent variables are extracted from clinical cardiac images according to a preset mode of corresponding image types, including:

[0089] When the image type is 3D image, the 3D spatial unwrapping block is used to process the clinical cardiac image to obtain the latent variables;

[0090] When the image type is a time series image, the temporal unwrapping block is used to process the clinical cardiac images to obtain latent variables;

[0091] When the image type is a 3D time series image, the spatial unwrapping block and the temporal unwrapping block are used together to process the clinical cardiac image to obtain the latent variables.

[0092] Taking the 3D images corresponding to computed tomography cardiac images and the 3D time series images corresponding to magnetic resonance cardiac images as examples, the anatomical encoder E in the 3D spatial unwrapping block ana1 The 128×128×128 sized 3D image block is processed and features are extracted through five 3D convolutional blocks containing skip connections and residual connections. After that, a probability map is generated through the attention mechanism to obtain the anatomical element N. At the same time, the modal encoder E in the 3D spatial disentanglement block mod1 The modal elements of the 3D image block are encoded to obtain the anatomical element M, and finally the 3D spatial disentanglement block outputs the latent variable.

[0093] Taking the time series images corresponding to ultrasound cardiac images and the 3D time series images corresponding to magnetic resonance images as examples, the anatomical encoder E in the time unwrapping block ana2 The time series images of length 10 are processed, first through multiple convolutional coding paths to avoid anatomical feature fusion, and then processed by the convolution block of ConvLSTM (convolutional long short-term memory network) structure to output the cache probability map to obtain the anatomical element N. At the same time, the modal encoder E in the time unwrapping block mod2 The modal elements of the time series images are encoded to obtain the anatomical elements M, and finally the time unwrapping block outputs the latent variables.

[0094] In one embodiment, a 3D spatial unwrapping block is used to process clinical cardiac images to obtain latent variables, including:

[0095] S31. Use the anatomical encoder in the 3D spatial unwrapping block to initially extract 3D anatomical features from the clinical cardiac image, obtaining a 3D anatomical feature map.

[0096] The CT image enters the anatomical encoder E of the 3D spatial unwrapping block as 3D image blocks of size 128×128×128. ana1 , and through the convolution operations of five 3D convolutional blocks, a 3D anatomical feature map is obtained. The 3D convolutional block contains multiple convolutional layers, and each convolutional layer is equipped with a specific convolutional kernel. The convolutional kernel slides on the three-dimensional space of the image to perform convolution operations on the image data. During the convolution operation, the parameters of the convolutional kernel are multiplied by the pixel values at the corresponding positions of the image and summed to extract local features of different scales and directions. To better capture features at different levels, the anatomical encoder also adopts skip connection and residual connection structures. The skip connection can directly transfer the feature information of the shallow layer to the deep layer, avoiding feature loss caused by vanishing gradients in the deep network, enabling the network to better learn the global features of the image. Further, through the convolution operations of five 3D convolutional blocks, a 3D anatomical feature map is obtained. Each convolutional block consists of two 3×3×3 convolutional layers, a batch normalization layer, and a PReLU, and skip connection and residual connection are adopted. There are also residual connections between the blocks, and the number of channels is 16, 32, 64, 128, 128 in sequence to perform initial 3D anatomical feature extraction, obtaining a 3D anatomical feature map.

[0097] S32. Use the attention mechanism to transform the 3D anatomical feature map into a probability map that matches the size of the clinical cardiac image, obtaining an intermediate result containing 3D anatomical features.

[0098] Taking the 3D anatomical feature map as the input, by calculating the importance weights of the features at each position in the entire image, focus on the regions that are significant for the cardiac anatomical structure. Specifically, through a series of operations such as linear transformation and activation functions, each feature vector in the 3D anatomical feature map is mapped to a probability value, and the probability value represents the importance degree of the feature at that position. Then, the probability value is weighted and fused with the 3D anatomical feature map, making the weights of the important features higher, thereby generating a probability map that matches the size of the clinical cardiac image. The probability map not only contains 3D anatomical features but also highlights the regions crucial for the recognition of the cardiac anatomical structure, and is an intermediate result containing 3D anatomical features.

[0099] S33. Use the modality encoder in the 3D spatial unwrapping block to perform modality extraction on the clinical cardiac image, obtaining an intermediate result containing CT image modality information.

[0100] Modality encoder E mod1The convolutional layer is used to extract the features of CT images, and the convolutional kernel parameters and network structure are designed according to the characteristics of CT images. CT images have specific gray-scale distributions, contrasts, and other features. The modality encoder captures these features through convolutional operations, thereby learning the unique modality information of CT images. Exemplarily, CT images show high contrast for bones and calcified tissues, and the modality encoder can extract the information related to CT imaging characteristics, forming an intermediate result containing the modality information of CT images.

[0101] S34. Integrate the intermediate result containing 3D anatomical features and the intermediate result containing the modality information of CT images to obtain a latent variable.

[0102] The integration method can be operations such as concatenation, weighted summation, etc. Exemplarily, in the case of concatenation, the two intermediate results are concatenated in the channel dimension to form a new feature vector, which contains 3D anatomical features and the modality information of CT images and constitutes the latent variable. Exemplarily, in the case of weighted summation, according to certain weight coefficients, the corresponding elements of the two intermediate results are weighted and added to obtain a latent variable containing both types of information.

[0103] In one embodiment, a temporal disentanglement block is used to process clinical cardiac images to obtain a latent variable, including:

[0104] S41. Use the anatomical encoder in the temporal disentanglement block to perform parallel processing on multiple convolutional encoding paths for the clinical cardiac image to obtain a probability map.

[0105] A time series of a set time series length is extracted from the US image and enters the anatomical encoder E in the temporal disentanglement block ana2 , the anatomical encoder E ana2 Performs parallel processing on the input image using multiple convolutional encoding paths. Each path independently extracts features from the image, aiming to capture the changing features of the cardiac anatomical structure in the time dimension from different angles. Each convolutional block consists of two 3×3×3 convolutional layers, a batch normalization layer, and a PReLU activation function. The first 3×3×3 convolutional layer performs a convolution operation on the input image, and the convolutional kernel slides in the three-dimensional space including the time dimension, and local features are extracted through the convolution operation. Exemplarily, it can capture the morphological change features of a local area of the heart at several adjacent time points. After the two convolutional layers, the batch normalization layer normalizes the convolved feature map. It calculates the mean μ and variance σ of the small batch data on each channel 2 , and then through the formula Normalize each element x, where γ and β are learnable parameters, and ∈ is a small constant to prevent the denominator from being zero. Batch normalization helps accelerate model convergence, reduce internal covariate shift, and make model training more stable. After batch normalization, the PReLU activation function introduces non-linearity. The expression of the PReLU function is where a is a learnable parameter, enhancing the non-linear expression ability of the model and enabling the model to learn more complex feature relationships.

[0106] To avoid anatomical feature fusion, skip connections and residual connections are adopted between each convolutional block, and there are also residual connections between blocks. The skip connection directly passes the shallow features to the deep layer to prevent information loss in the deep network; the residual connection allows the network to learn the residual between the input and the output. Exemplarily, as the convolutional encoding path progresses, the number of channels is 16, 32, 64, 128, 128 in sequence. Such a setting enables the model to gradually learn richer and more representative anatomical features. After parallel processing through four convolutional encoding paths, the feature maps generated by each path are concatenated in the channel dimension to form a comprehensive feature map. Then, a 1×1×1 convolutional layer is used to process the comprehensive feature map to adjust the number of channels to meet the requirements of subsequent probability map generation. Finally, the Softmax function is used to convert the comprehensive feature map into a probability map with a size of 128×128, and the value of each pixel represents the probability that the position belongs to different anatomical structures.

[0107] S42. Use a convolutional block to perform convolutional processing on the probability map to convert it into a cached probability map, generating an intermediate result of anatomical elements containing temporal features.

[0108] The convolutional block used for the probability map is also composed of two 3×3×3 convolutional layers, a batch normalization layer, and a PReLU activation function. After processing by the convolutional block, a cached probability map is obtained. Compared with the probability map, the cached probability map has more abstract and refined features and can better reflect the core features of the cardiac anatomical structure changing over time. Further, a global average pooling operation is performed on the cached probability map to compress the spatial dimension to 1, obtaining a vector containing only temporal features. Then, the dimension is adjusted through a fully connected layer to generate an intermediate result of anatomical elements containing temporal features. The intermediate result contains key feature information of the cardiac anatomical structure in the temporal dimension, such as the size changes of each cardiac chamber at different time points and the temporal pattern of myocardial movement.

[0109] S43. Use the modality encoder in the temporal disentanglement block to perform modality extraction on clinical cardiac images to obtain an intermediate result containing time series modality information.

[0110] Modality encoder E mod2It is composed of multiple convolutional layers, batch normalization layers, and activation functions, and is specifically designed for extracting modal features of time series images. Schematically, a series of 3×3×3 convolutional layers are used to perform convolutional operations on the input image to capture spatio-temporal features at different scales. Parameters such as the number of convolutional kernels and the stride of the convolutional layer will be adjusted according to the specific task and data characteristics. Exemplarily, the first convolutional layer can use fewer 8 convolutional kernels to initially extract some basic modal features. As the network deepens, the number of convolutional kernels gradually increases to 16, 32, etc. to learn more complex modal features. To better capture the time series modal information, convolutional operations in the time dimension may be introduced. Exemplarily, 1D convolution extracts features on the time axis. The 1D convolutional kernel slides in the time dimension to integrate the image features of different time frames and extract the key information of the time series modality. Specifically, features related to the modality such as the change pattern of the echo intensity in the time series of ultrasonic images can be captured through 1D convolution.

[0111] S44. Integrate the intermediate result of the anatomical element containing time features and the intermediate result containing time series modal information to obtain a latent variable.

[0112] An integration method of concatenation in the channel dimension can be adopted. Exemplarily, the dimension of the intermediate result of the anatomical element containing time features is D1, and the dimension of the intermediate result containing time series modal information is D2. The dimension of the obtained latent variable after concatenation is D1 + D2. Through concatenation, the latent variable contains both the change information of the cardiac anatomical structure in the time dimension and the specific information of the US time series imaging modality, providing a comprehensive and targeted information basis for subsequent tasks such as image reconstruction and segmentation based on this latent variable.

[0113] In one embodiment, a spatial unwrapping block and a temporal unwrapping block are used together to process clinical cardiac images to obtain a latent variable, including:

[0114] S51. Use the anatomical encoder in the 3D spatial unwrapping block to initially extract 3D anatomical features from the clinical cardiac image to obtain a 3D anatomical feature map.

[0115] The MR image enters the anatomical encoder E of the 3D spatial unwrapping block in the form of a 3D image block with a size of 128×128×128 ana1 , and a 3D anatomical feature map is obtained through the convolutional operations of five 3D convolutional blocks.

[0116] S52. Use the attention mechanism to transform the 3D anatomical feature map into a probability map that matches the size of the clinical cardiac image to obtain an intermediate result containing 3D anatomical features.

[0117] The 3D anatomical feature map enters the attention mechanism to generate a probability map that matches the size of the input image, obtaining an intermediate result containing 3D anatomical features.

[0118] S53. Use the modality encoder in the 3D spatial unwrapping block to extract the modality of the clinical cardiac image, obtaining an intermediate result containing the modality information of the MR image.

[0119] Modality encoder E mod1 Process the MR image, and finally output an intermediate result containing the modality information of the MR image through the fully connected layer.

[0120] S54. Use the anatomical encoder in the temporal unwrapping block to perform parallel processing on multiple convolutional encoding paths of the clinical cardiac image to obtain a probability map.

[0121] Extract a time series of length 10 from the MR data and enter the anatomical encoder E of the temporal unwrapping block. ana2 . Perform parallel processing on multiple convolutional encoding paths, each path processes an image of 128×128, avoiding the fusion of anatomical features, generating a probability map, and then adding a dropout of 0.2 to prevent overlap.

[0122] S55. Use the convolutional block to perform convolutional processing on the probability map to convert it into a cached probability map, generating an intermediate result of anatomical elements containing temporal features.

[0123] Use four convolutional blocks with ConvLSTM structures to process the probability map. The input is the combination of probability maps of the entire sequence, and the output is the cached probability map, thereby obtaining an intermediate result of anatomical elements containing temporal features.

[0124] S56. Use the modality encoder in the temporal unwrapping block to extract the modality of the clinical cardiac image, obtaining an intermediate result containing the modality information of the time series.

[0125] Modality encoder E mod2 Similarly process the MR time series image, and finally output an intermediate result containing the modality information of the time series through the fully connected layer.

[0126] S57. Fuse the intermediate result containing 3D anatomical features and the intermediate result of anatomical elements containing temporal features to obtain anatomical elements.

[0127] Fuse the intermediate results of anatomical elements obtained by the 3D spatial unwrapping block and the temporal unwrapping block respectively, integrating the 3D spatial features and temporal features to form the final anatomical element N.

[0128] S58. Fuse the intermediate result containing the modality information of the MR image and the intermediate result containing the modality information of the time series to obtain modality elements.

[0129] E mod1and E mod2 The output modal information is also integrated to obtain the final modal element M.

[0130] S59. The anatomical element and the modal element constitute latent variables.

[0131] The anatomical element N and the modal element M will be used as the output of the dynamic causal module, that is, latent variables, and will be sent to subsequent diffusion modules for further processing.

[0132] In one embodiment, as Figure 3 shown, a cardiac reconstruction image is generated according to the latent variable, including:

[0133] S301. Generate an initial image using the latent variable;

[0134] The anatomical element determines the general shape, position, and dynamic changes of the cardiac anatomical structure in the initial image. Exemplarily, the time pattern of cardiac systole and diastole encoded in the anatomical element will guide the generated initial image to present corresponding cardiac morphological changes in the time series. The modal element affects the imaging characteristics of the initial image, making it conform to the characteristics of ultrasound images. Exemplarily, the ultrasonic echo characteristic information in the modal element will make the generated initial image similar to a real ultrasound image in terms of gray-scale distribution, and the boundaries and textures of different tissues have the characteristics of ultrasound images. Through the step-by-step decoding and mapping of the latent variable by the decoder, the finally generated initial image reflects to a certain extent the cardiac anatomical structure and ultrasonic imaging characteristics, providing a basic template for the subsequent generation of the reconstruction image, and further optimizing it through the diffusion process to make it closer to the real cardiac ultrasound image.

[0135] S302. Add Gaussian noise to the initial image to obtain a noisy image.

[0136] According to the preset variance schedule β1, β2, … β r , Gaussian noise is gradually added to the initial image. Each addition of noise is based on the image state of the previous step, and the formula is expressed as:

[0137]

[0138] where r represents the step of adding noise, from 1 to R, is the noisy image after the r-th addition of noise.

[0139] As the noise is gradually added, the image quality gradually deteriorates. After R steps of operation, an image sequence with different noises is obtained This process simulates the change of data from clear to blurred, providing diverse noisy images for the reconstruction of the inverse sub-module.

[0140] S303. Use Gaussian transformation to reshape the noisy image to obtain a cardiac reconstruction image.

[0141] Starting from the image with added noise as the starting point, reconstruct the image through the learned Gaussian transformation, and its Gaussian transformation process is expressed as:

[0142]

[0143] Start from r = R and gradually reconstruct forward to finally obtain the cardiac reconstruction image It is difficult to directly obtain the reverse distribution during the reconstruction process but it can be obtained using Bayes' formula to ensure the similarity of the cardiac reconstruction image with x0 of the original clinical cardiac image. During the diffusion processes of different modalities, there are differences in the reverse processes, and the network generates corresponding cardiac reconstruction images according to the specific generation of different modalities.

[0144] In the second aspect, as Figure 4 shown, the present application also provides a spatio-temporal causal cross-modal cardiac image segmentation system, which includes:

[0145] An image acquisition module for acquiring clinical cardiac images.

[0146] An image recognition module for identifying the image type of the clinical cardiac image and obtaining the annotation information corresponding to the image type.

[0147] A model algorithm module for acquiring the clinical cardiac image and the annotation information, and performing image segmentation according to different image types in a set pattern to obtain the cardiac image segmentation result.

[0148] Construction of a spatio-temporal causal diffusion network. Define the set of clinical cardiac images as X, and the corresponding annotation data as Y. Introduce latent variables, namely the modality element M and the anatomical element N, to simulate the generation processes of observing X and Y. Among them, the latent variable M represents the modality element of the cardiac image, such as the cardiac background, and N represents the anatomical element, which usually includes the shape and size of the heart. Further, the time series features, 3D features, and 3D time features are also included in the anatomical element. The cardiac reconstruction image is generated by the combination of M and N. As long as the intervention of the anatomical element N causes a change in the label Y, learn invariant prediction by introducing the evidence lower bound (ELBO), so as to learn invariant representations, reconstruct past images, and minimize the distance between the prediction and the true value.

[0149] In one embodiment, as Figure 5 shown, the model algorithm module includes a dynamic causal module, a diffusion module, and a segmentation module;

[0150] The dynamic causal module is used to extract latent variables from clinical cardiac images in accordance with a preset pattern corresponding to the image type by means of causal intervention; causal intervention is that the dynamic causal module utilizes the characteristics of the dynamic causal module objective function reflecting the spatio-temporal causal relationship between the cardiac segmentation image and the annotation data, and extracts latent variables to determine the characteristics of the spatio-temporal causal relationship between the cardiac segmentation image and the annotation data;

[0151] The dynamic causal module extracts causally invariant representations spatio-temporally by means of causal intervention. For clinical cardiac images of different image types, different objective functions are defined respectively. The objective function depends on previous interventions, and latent variable extraction guidance is obtained by exploring the causal relationship between the objective functions at consecutive time steps, so as to consider the causal relationship between the input clinical cardiac image and the output cardiac image segmentation structure and the past optimal intervention, promote information fusion, and accelerate the process of determining the current optimal intervention. Further, the dynamic causal module designs 3D spatial disentangling blocks and temporal disentangling blocks according to the data characteristics of different image types, which include different encoders to extract latent variables.

[0152] The diffusion module is used to generate cardiac reconstruction images according to the latent variables.

[0153] The cardiac reconstruction image is synthesized by optimizing logp(x|m,n) to be closer to the real image. The diffusion module includes forward and backward processes. The forward process gradually adds Gaussian noise to the samples, and the backward process generates samples based on the learned Gaussian transfer. The relevant distribution is obtained through Bayes' formula to ensure the similarity between the synthesized cardiac reconstruction image and the original clinical cardiac image.

[0154] During the model construction process, during the training process of the diffusion module, logp(x|m,n) is continuously optimized, and the feature representations of M and N are further adjusted through the forward and backward processes, changing the influence of M and N on image generation, synthesizing the cardiac reconstruction image, and ensuring that the anatomical features represented by N are causally invariant. The cardiac reconstruction image is obtained After that, the reconstruction loss is calculated By comparing the difference between the reconstructed image and the original clinical cardiac image x0 to evaluate the reconstruction effect, the loss value is fed back to the training process of the diffusion module to adjust the parameters in the diffusion module, optimize the forward and backward processes, make the reconstructed image closer to the original image, and further improve the model's processing ability for anatomical elements and modal elements, and enhance the performance of the spatio-temporal causal-based cross-modal cardiac image segmentation system.

[0155] The segmentation module is used to perform convolutional processing on the cardiac reconstruction image to obtain the cardiac image segmentation result.

[0156] Based on the anatomical element N, different convolutional modules are used to predict the annotation information according to the modality of the cardiac reconstruction image. By defining the loss function LS to optimize the prediction results.

[0157] In one embodiment, when the image type is a 3D image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship features between the cardiac segmentation image and the annotation data:

[0158] argminE[Y S |do(X s = x s ), I 0:S-1

[0159]

[0160] where argmin is the intervention strategy for minimizing the expected loss, E[Y S |do(X s = x s ), I 0:S-1 is the expected condition; do)X s = x s ) is the causal intervention, I 0:S-1 is all the interventions and the expected values of the corresponding labels Y S in the known previous spatial dimensions; Y S is the annotation data;

[0161] When the image type is a time series image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship features between the cardiac segmentation image and the annotation data:

[0162] argminE[Y t |do(X t = x t ,), I 0:t-1

[0163]

[0164] where argmin is the intervention strategy for minimizing the expected loss, E[Y t |do(X t = x t ,), I 0:t-1 is the expected condition; do(X t = x t ,) is the causal intervention, I 0:t-1 is all the interventions and the expected values of the corresponding labels Y t in the known previous time dimensions; Y t is the annotation data;

[0165] ​​When the image type is 3D time series images, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship characteristics between the cardiac segmentation images and the annotation data:

[0166] argminE[Y t,s |do(X t.s =x t,s ),I 0:t-1,s-1

[0167]

[0168] where argmin is the intervention strategy that minimizes the expected loss, and E[Y t,s |do(X t.s =x t,s ),I 0:t-1,s-1 is the expected condition; do(X t.s =x t,s ) is the causal intervention, and I 0:t-1,s-1 is all the interventions I 0:t-1 in the previous time dimension and all the interventions I 0:S-1 in the previous spatial dimension, as well as the expected value of the jointly corresponding label Y t,s ; Y t,s is the annotation data.

[0169] The cross-modal cardiac image segmentation system based on spatio-temporal causality provided in this embodiment, in the process of constructing the spatio-temporal causal diffusion network model, according to the causal graph and causal intervention, combines the intervention information at different time steps, and continuously adjusts the learning of the anatomical element N and the modal element M. For example, for time series images, through argminE[Y t |do(X t =x t ,),I 0:t-1 , the objective function is calculated, and the previous interventions are fully considered for the causal relationship in the time dimension. By continuously exploring the causal relationship of the objective function at consecutive time steps, the learning of the latent variables can be optimized, and then the causal features related to the annotation information Y can be better extracted, promoting information fusion and accelerating the process of determining the optimal intervention. The objective function is used to guide the extraction of latent variables. However, in the training stage, the latent variables provide the calculation basis for the objective function, and thus affect the adjustment and optimization of the model parameters.

[0170] ​It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown sequentially according to the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0171] The above-described embodiments only represent several implementation manners of the embodiments of the present application, and the description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the embodiments of the present application.

Claims

1. A cross-modal cardiac image segmentation method based on spatio-temporal causality, characterized in that, The method includes: Obtaining a clinical heart image; the clinical heart image includes a magnetic resonance heart image, a computed tomography heart image, and an ultrasonic heart image; Identifying the image type of the clinical heart image and obtaining annotation information corresponding to the image type; the image types include time series images, 3D images, and 3D time series images; Inputting the clinical heart image and the annotation information into a spatio-temporal causal diffusion network model to obtain a heart image segmentation result; the spatio-temporal causal diffusion network model is used to perform image segmentation on the clinical heart image according to the corresponding image type in a set mode to obtain a heart image segmentation result; the heart image segmentation result includes a heart segmentation image and annotation data.

2. The method according to claim 1, characterized in that, The performing image segmentation on the clinical heart image according to the corresponding image type in a set mode to obtain a heart image segmentation result includes: Using causal intervention to extract latent variables from the clinical heart image according to the preset mode corresponding to the image type; the latent variables include anatomical elements and modal elements; Generating a heart reconstruction image according to the latent variables; Performing convolution processing on the heart reconstruction image and the anatomical elements to obtain a heart image segmentation result.

3. The method according to claim 2, wherein The extracting latent variables from the clinical heart image according to the preset mode corresponding to the image type includes: When the image type is a 3D image, using a 3D spatial unwrapping block to process the clinical heart image to obtain the latent variables; When the image type is a time series image, using a time unwrapping block to process the clinical heart image to obtain the latent variables; When the image type is a 3D time series image, using a spatial unwrapping block and a time unwrapping block to jointly process the clinical heart image to obtain the latent variables.

4. The method according to claim 3, wherein The using a 3D spatial unwrapping block to process the clinical heart image to obtain the latent variables includes: Using an anatomical encoder in the 3D spatial unwrapping block to preliminarily extract 3D anatomical features from the clinical heart image to obtain a 3D anatomical feature map; Using an attention mechanism to convert the 3D anatomical feature map into a probability map matching the size of the clinical heart image to obtain an intermediate result containing 3D anatomical features; Using a modal encoder in the 3D spatial unwrapping block to perform modal extraction on the clinical heart image to obtain an intermediate result containing CT image modal information; Integrating the intermediate result containing 3D anatomical features and the intermediate result containing CT image modal information to obtain latent variables.

5. The method according to claim 3, characterized in that The using a time unwrapping block to process the clinical heart image to obtain the latent variables includes: Using an anatomical encoder in the time unwrapping block to perform parallel processing on multiple convolutional encoding paths of the clinical heart image to obtain a probability map; Using a convolutional block to perform convolution processing on the probability map to convert it into a cached probability map and generating an intermediate result of anatomical elements containing time features; Using a modal encoder in the time unwrapping block to perform modal extraction on the clinical heart image to obtain an intermediate result containing time series modal information; Integrate the intermediate result of the anatomical element containing temporal features and the intermediate result of the intermediate result containing time series modality information to obtain a latent variable.

6. The method according to claim 3, wherein The use of the spatial unwrapping block and the temporal unwrapping block together to process the clinical heart image to obtain the latent variable includes: Use the anatomical encoder in the 3D spatial unwrapping block to preliminarily extract 3D anatomical features from the clinical heart image to obtain a 3D anatomical feature map; Use the attention mechanism to convert the 3D anatomical feature map into a probability map matching the size of the clinical heart image to obtain an intermediate result containing 3D anatomical features; Use the modality encoder in the 3D spatial unwrapping block to perform modality extraction on the clinical heart image to obtain an intermediate result containing MR image modality information; Use the anatomical encoder in the temporal unwrapping block to perform parallel processing of multiple convolutional encoding paths on the clinical heart image to obtain a probability map; Use the convolutional block to perform convolutional processing on the probability map to convert it into a cached probability map, generating an intermediate result of the anatomical element containing temporal features; Use the modality encoder in the temporal unwrapping block to perform modality extraction on the clinical heart image to obtain an intermediate result containing time series modality information; Fuse the intermediate result containing 3D anatomical features and the intermediate result of the anatomical element containing temporal features to obtain the anatomical element; Fuse the intermediate result containing MR image modality information and the intermediate result containing time series modality information to obtain the modality element; The anatomical element and the modality element constitute the latent variable.

7. The method according to claim 2, characterized in that, The generation of the cardiac reconstruction image according to the latent variable includes: Generate an initial image using the latent variable; Add Gaussian noise to the initial image to obtain a noisy image; Use Gaussian transfer to reshape the noisy image to obtain a cardiac reconstruction image.

8. A cross-modal cardiac image segmentation system based on spatio-temporal causality, characterized in that, The system includes: An image acquisition module for acquiring clinical heart images; An image recognition module for identifying the image type of the clinical heart image and obtaining the annotation information corresponding to the image type; A model algorithm module for acquiring the clinical heart image and the annotation information, and performing image segmentation on the clinical heart image according to different image types in a set mode to obtain a cardiac image segmentation result.

9. The system according to claim 8, wherein: The model algorithm module includes a dynamic causality module, a diffusion module, and a segmentation module; The dynamic causality module is used to extract a latent variable from the clinical heart image according to a preset mode corresponding to the image type by using causal intervention; the causal intervention is that the dynamic causality module uses the objective function of the dynamic causality module to reflect the spatio-temporal causal relationship characteristics of the cardiac segmentation image and the annotation data, and extracts the latent variable to determine the spatio-temporal causal relationship characteristics of the cardiac segmentation image and the annotation data; The diffusion module is used to generate a cardiac reconstruction image according to the latent variable; The segmentation module is used to perform convolutional processing on the cardiac reconstruction image to obtain a cardiac image segmentation result.

10. The system according to claim 9, wherein: When the image type is a 3D image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship features between the cardiac segmentation image and the annotation data: argminE[Y S |do(X s =x s ),I 0:S-1 ​ Among them, argmin is the intervention strategy that minimizes the expected loss, and E[Y S |do(X s =x s ),I 0:S-1 is the expected condition; do(X s =x s ) is the causal intervention, and I 0:S-1 is all the interventions of the known previous spatial dimensions and the expected value of the corresponding label Y S ; Y S is the labeled data; When the image type is a time series image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship features between the cardiac segmentation image and the annotation data: argminE[Y t |do(X t =x t ,),I 0:t-1 ​ where argmin is the intervention strategy that minimizes the expected loss, E[Y t |do(X t = x t ,), I 0:t-1 is the expected condition; do(X t = x t ,) is the causal intervention, I 0:t-1 is the expected value of all interventions and the corresponding labels Y t in the known previous time dimension; Y t is the labeled data; When the image type is a 3D time series image, the following dynamic causal module objective function is used to reflect the spatio-temporal causal relationship features between the cardiac segmentation image and the annotation data: argminE[Y t,s |do(X t.s =x t,s ),I 0:t-1,s-1 Among them, argmin is the intervention strategy that minimizes the expected loss, E[Y t,s |do(X t.s = x t,s 0, I 0:t-1,s-1 is the expected condition; do(X t.s = x t,s ) is a causal intervention, I 0:t-1,s-1 is all interventions I in the known previous time dimension 0:t-1 and all interventions I in the previous space dimension 0:S-1 and the corresponding label Y t,s of the expectation value; Y t,s is the labeled data.

Citation Information

Cited By

  • Cancer prediction model construction method based on causal network and adaptive feature selection

    CN122050850A

  • Cancer prediction model construction method based on causal network and adaptive feature selection

    CN122050850B