Three-dimensional dynamic medical image diagnosis method and system based on self-supervised multi-modal time sequence fusion, electronic equipment and storage medium
By employing a self-supervised multimodal temporal fusion method, utilizing modal confidence attention and phase-aware feature encoders, the problems of insufficient generalization ability and poor robustness of multimodal fusion in cross-center image analysis are solved, achieving efficient and interpretable medical image diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-03-27
AI Technical Summary
Existing medical image analysis methods lack generalization ability under cross-center and cross-device conditions, have poor robustness in multimodal information fusion, rely on manual annotation at high cost, and lack self-supervised learning and clinical interpretability.
A self-supervised multimodal temporal fusion method is adopted, which combines a modal confidence attention fusion mechanism and a phase-aware feature encoder with 3D mask prediction and phase sequence comparison learning to achieve adaptive weighted fusion of multimodal images and interpretability of diagnostic results.
It improves the accuracy and robustness of cross-center diagnostics, reduces reliance on manual annotation, provides interpretable diagnostic evidence, and achieves an efficient end-to-end diagnostic process.
Smart Images

Figure CN121746790A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of medical image processing and artificial intelligence, in particular, especially relates to a three-dimensional dynamic medical image diagnosis method and system based on self-supervised multi-modal time sequence fusion, an electronic device and a storage medium. The present application can be directly applied to computer-aided diagnosis in clinical fields such as oncology, cardiology, and neuroimaging, and is especially suitable for lesion segmentation, disease classification, efficacy evaluation, and prognosis prediction based on dynamic contrast-enhanced images obtained by magnetic resonance imaging (MRI), computed tomography (CT), and other devices. BACKGROUND
[0002] With the wide application of deep learning technology in medical image analysis, artificial intelligence-aided diagnosis has become one of the core driving forces for the development of clinical imaging. In key fields such as oncology, cardiology, and neuroimaging, computer vision-based automated image processing methods, such as algorithms based on convolutional neural networks (CNN), visual Transformers (ViT), and their hybrid architectures, have been widely used for lesion segmentation, disease classification, efficacy evaluation, and prognosis prediction in clinical tasks. In particular, the development of three-dimensional imaging and dynamic contrast-enhanced imaging technology (such as dynamic contrast-enhanced magnetic resonance imaging DCE-MRI, dynamic contrast-enhanced CT, etc.) enables clinicians to observe the dynamic changes in structure and function of organs and lesions, providing unprecedented rich information for precise diagnosis and personalized treatment of diseases.
[0003] In the prior art, some algorithms have been able to achieve high-precision automatic segmentation and classification in single-modal three-dimensional medical images. In view of the time sequence characteristics of dynamic contrast-enhanced images, researchers have introduced time convolution networks (TCN) or time attention mechanisms to capture time-varying features such as blood flow dynamics of lesions during contrast agent perfusion. In addition, self-supervised learning (Self-Supervised Learning, SSL) strategies have also been applied to the field of medical images, through proxy tasks such as mask reconstruction, contrast learning, etc. without the need for manual annotation, learning universal feature representations from massive unlabeled data, thereby reducing the dependence on expensive manual annotation to some extent and improving the generalization ability of the model.
[0004] However, despite the above progress, existing technologies still face a series of severe and fundamental challenges in actual clinical scenarios: Insufficient generalization capability: The problem of distribution drift across medical centers, imaging devices, and scanning protocols is significant. Differences in imaging devices (e.g., MRI or CT machines from different manufacturers or models), scanning parameters (e.g., TR, TE, slice thickness), and contrast agent usage protocols among different hospitals result in systematic deviations in image gray scale distribution, spatial resolution, signal-to-noise ratio, and artifact patterns. Existing models often struggle to be directly applied to data from other centers after training on a single dataset, resulting in a sharp decline in performance.
[0005] Annotation data bottleneck: The cost of obtaining high-quality annotated data is extremely high. In the context of three-dimensional and four-dimensional (three-dimensional + time) dynamic images, precise annotation by radiologists layer by layer and phase by phase is not only time-consuming and labor-intensive, but also highly susceptible to the experience level and subjective judgment of the annotators, leading to label noise and inconsistent annotation quality, which directly limits the performance ceiling and reliability of supervised learning models.
[0006] Insufficient and costly modeling of temporal information: Existing methods rely heavily on explicit 4D convolution or global time series Transformer architectures to handle dynamic images. While these methods can theoretically capture rich temporal dependencies, they suffer from excessive memory usage, high computational complexity, and long inference delays when dealing with high-resolution three-dimensional medical images, making them unsuitable for efficient deployment and real-time application on servers or image workstations (PACS) in clinical environments.
[0007] Poor robustness of multi-modal fusion: Clinical diagnosis often requires integrating information from multiple modalities of images (e.g., T1, T2, DWI sequences of MRI). Existing fusion methods (e.g., simple concatenation or weighted averaging) often assume that the data quality of all modalities is reliable. However, in practice, a modality may be of poor quality or even unavailable due to patient movement, equipment failure, or scanning artifacts. The lack of a mechanism to perceive modality quality or confidence makes the model significantly less robust when faced with low-quality or abnormal modalities, and it may even be misled by noisy modalities.
[0008] Lack of clinical explainability and trustworthiness: Most deep learning models are like "black boxes," and their decision-making process is difficult for clinicians to understand and trace. In critical decision-making scenarios such as rare diseases, long-tail clinical cases, or critical conditions, the lack of explainable diagnostic evidence can severely limit the adoption of algorithms in actual work and the trust of doctors.
[0009] The root cause of these problems is that existing medical image algorithms mostly rely on single modality, static images and supervised labels, lack a self-supervised modeling strategy that can simultaneously and efficiently utilize multi-modal information and dynamic temporal characteristics, and have not formed a mechanism that is highly consistent with clinical needs and can perceive data quality and perform adaptive information fusion.
[0010] In summary, there is an urgent need in the art for a new technical solution that can efficiently model the temporal information of dynamic images, robustly fuse multi-modal features, and reduce the dependence on manual annotation through self-supervised learning, thereby effectively improving the generalization performance of the model under cross-center and cross-device conditions, the diagnostic accuracy for rare diseases, and the stability and interpretability of dynamic analysis, without significantly increasing the computational burden. SUMMARY
[0011] The present application aims to solve the problems in the prior art, such as the insufficient generalization ability of three-dimensional and dynamic enhancement medical image analysis methods when processing cross-center data, the insufficient modeling of dynamic temporal characteristics, the poor robustness of multi-modal information fusion, the strong dependence on high-quality manual annotation, and the low diagnostic accuracy for rare diseases. The present application proposes a three-dimensional dynamic medical image diagnosis method, system, electronic device and storage medium based on self-supervised multi-modal temporal fusion to solve the above technical problems.
[0012] The first aspect of the present application discloses a three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal temporal fusion, the method comprising: obtaining one or more modal three-dimensional dynamic medical image sequences, the sequence containing multiple phases arranged in chronological order; standardizing and preprocessing the image sequence to generate a standardized four-dimensional tensor; using a phase-aware feature encoder to extract multi-scale spatio-temporal features for each modality from the four-dimensional tensor; using a modality confidence attention fusion mechanism to adaptively weight and fuse the multi-scale spatio-temporal features according to the confidence estimated in real time for each modality, to generate fusion features; and inputting the fusion features into one or more downstream task heads to generate diagnosis results; wherein the parameters of the feature encoder and the fusion mechanism are determined through a training process that includes a self-supervised optimization task, and the self-supervised optimization task includes: a three-dimensional mask reconstruction task based on partially occluding regions of input images or features and predicting the contents of the occluded regions; and a phase order contrast learning task based on distinguishing correct phase order and incorrect phase order.
[0013] According to the method of the first aspect of the present application, the standardization preprocessing step comprises: spatially registering images of different modalities or phases to align anatomical structures; resampling images to isotropic voxel resolution; intensity normalizing images to reduce intensity distribution differences introduced by different devices or scanning protocols; and unifying the number of temporal phases of all image sequences to a preset length by interpolation or downsampling.
[0014] According to the method of the first aspect of the present application, the step of the phase-aware feature encoder extracting multi-scale spatio-temporal features comprises: dividing the volume data of each phase into three-dimensional voxel blocks, and embedding a phase position code and a spatial position code for each voxel block; constructing and fusing the difference features or ratio features of the voxel blocks between adjacent phases to enhance the representation of temporal dynamic changes; and processing the voxel block embeddings using a stacked multi-head self-attention module, wherein the self-attention adopts a windowing strategy in the temporal dimension.
[0015] According to the method of the first aspect of the present application, the modality confidence attention fusion mechanism comprises: for each modality, using a quality assessment subnetwork to estimate a local confidence map reflecting the local data quality and a global confidence score reflecting the overall data quality from its spatio-temporal features; based on the local confidence map and the global confidence score, calculating the fusion weight of each modality at each spatio-temporal position; and weighting and summing the features of each modality according to the fusion weight to obtain the fused features.
[0016] According to the method of the first aspect of the present application, the training process further comprises stability and generalization constraints, the constraints being selected from at least one of the following: deformation equivariance constraint, which requires the model to maintain equivariance in the feature space for geometric deformation applied to the input image; confidence consistency constraint, which requires the model to maintain consistency in the estimated modality confidence for different views or perturbed versions of the same image.
[0017] According to the method of the first aspect of the present application, the downstream task head comprises at least one of a segmentation head, a classification head, a detection head or a regression head; the diagnostic result comprises a segmentation mask of a lesion, a classification probability of a disease, a detection bounding box of a lesion, or a voxel-level physiological parameter map.
[0018] The method according to the first aspect of the present application further comprises a step of post-processing the diagnostic result, the post-processing comprising: temperature scaling calibration of the classification probability; temporal smoothing of the segmentation result or parameter map of the dynamic sequence; and packaging and outputting the diagnostic result and / or the modality confidence map in a format conforming to the DICOM standard.
[0019] The second aspect of the present application discloses a three-dimensional dynamic medical image diagnosis system based on self-supervised multi-modal temporal fusion, which adopts the method according to any one of the first aspect, and the system comprises: a data preprocessing module configured to acquire and process a three-dimensional dynamic medical image sequence of one or more modalities to generate a standardized four-dimensional tensor; a feature encoding module configured to extract multi-scale spatio-temporal features for each modality from the four-dimensional tensor by using a phase-aware feature encoder; a multi-modal fusion module configured to perform adaptive weighted fusion on the multi-scale spatio-temporal features according to a confidence of each modality estimated in real time by a modality confidence attention fusion mechanism to generate fusion features; a diagnosis reasoning module configured to input the fusion features into one or more downstream task heads to generate a diagnostic result; wherein the parameters of the feature encoding module and the multi-modal fusion module are determined through a training process comprising a self-supervised optimization task, and the self-supervised optimization task module is configured to perform a self-supervised optimization task, the self-supervised optimization task comprising at least one of: a three-dimensional mask reconstruction task of predicting the content of a blocked area by blocking a partial area of an input image or feature; a phase sequence contrast learning task based on distinguishing correct phase sequences from incorrect phase sequences.
[0020] The third aspect of the present application discloses an electronic device. The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal temporal fusion according to any one of the first aspect of the present application.
[0021] The fourth aspect of the present application discloses a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal temporal fusion according to any one of the first aspect of the present application.
[0022] Compared with the prior art, the present application has the following remarkable advantages and beneficial effects: 1. Significantly improve diagnostic accuracy and dynamic analysis capability: By combining three-dimensional mask prediction and phase sequence contrast learning, the present application can effectively learn the fine spatial structure and subtle temporal evolution characteristics of lesions from unlabeled data. This makes the model perform better in dynamic perfusion analysis, early lesion detection and disease evolution tracking, thereby improving the sensitivity and specificity of diagnosis.
[0023] 2. Enhance cross-center generalization ability and robustness: The present application uses deformation invariance and confidence consistency as generalization constraints to make the model more resistant to data distribution drift caused by cross-center, cross-device and cross-protocol. At the same time, the modal confidence attention (MCA) mechanism can adaptively reduce the influence when the quality of a certain modality is poor or missing, ensuring the robustness of multi-modal fusion and the stability of diagnostic results.
[0024] 3. Greatly reduce the dependence on manual annotation and cost: The present application takes self-supervised learning as the core, which can use a large amount of unlabeled or weakly labeled clinical data for model pre-training. This greatly reduces the dependence on expensive and time-consuming pixel-level manual annotation, reduces the threshold and cost of data preparation, and makes it possible to build high-performance models.
[0025] 4. Provide interpretable diagnostic basis and enhance clinical trust: The MCA mechanism can explicitly output the contribution weight distribution of each modality in the diagnostic decision, and intuitively show the doctor which information sources the model "trusts". Combined with the visualization of phase difference / ratio features, it can quantitatively reflect the temporal dynamics of lesions, providing objective and traceable quantitative basis for doctors' diagnosis, staging and efficacy evaluation, and solving the trust crisis of "black box" models.
[0026] 5. Realize efficient end-to-end diagnosis process: The overall framework proposed by the present application constitutes an end-to-end automated process from data input to multi-task (segmentation, classification, regression) result output. The windowed temporal attention and lightweight fusion design used in the present application control the computational overhead while ensuring performance, making efficient inference on standard clinical servers possible and having good prospects for engineering deployment. BRIEF DESCRIPTION OF DRAWINGS
[0027] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor. Figure 1 is a flowchart of a three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal temporal fusion according to an embodiment of the present application. Figure 2 is a self-supervised optimization task structure diagram according to an embodiment of the present application; Figure 3 is a modal confidence attention fusion mechanism diagram according to an embodiment of the present application; Figure 4 is a structure diagram of a three-dimensional dynamic medical image diagnosis system based on self-supervised multi-modal temporal fusion according to an embodiment of the present application; Figure 5 is a structure diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the objects, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0029] Embodiment 1: The three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal temporal fusion proposed in the present application aims to solve the problems of insufficient cross-center generalization, insufficient dynamic temporal modeling and poor multi-modal fusion robustness in the prior art. The overall idea of the method is as follows: taking multi-modal dynamic three-dimensional images as input, the images are preprocessed and standardized to construct a four-dimensional tensor representation capable of expressing spatial structure and time sequence information; on this basis, combined with phase perception feature modeling and modal confidence weighting mechanism, multi-source image features are extracted and fused; then the model is trained through self-supervised tasks and stability constraints, and finally the unified output of diagnosis tasks such as lesion segmentation, disease classification and perfusion parameter regression is realized.
[0030] In the method of the present application, the data acquisition and preprocessing link is the basic link of the whole technical solution, and its main task is to convert the multi-modal three-dimensional dynamic medical images from different devices, different centers and different protocol conditions into unified, standardized and comparable input data, so as to facilitate the subsequent feature modeling and fusion. In actual clinical application, the differences of image data are manifested as differences in device manufacturers, inconsistent scanning parameters, diversified use of enhancers, different spatial resolutions and layer thicknesses, and differences in motion control of different patients. If not properly preprocessed, these differences will directly lead to the distribution deviation of the subsequent network input data, and then seriously affect the generalization performance and diagnostic stability of the model. Therefore, this section explains from five aspects of input data collection and format unification, spatial registration and resampling, intensity normalization and contrast adjustment, time phase ordering and alignment, and data quality control.
[0031] The first aspect of the present application discloses a three-dimensional dynamic medical image diagnosis method based on self-supervised multi-modal time sequence fusion. As shown in the figure, Figure 1 The method comprises the following steps: Step S1: acquiring a three-dimensional dynamic medical image sequence of one or more modalities, the sequence containing a plurality of phases arranged in time sequence; performing standardized preprocessing on the image sequence to generate a standardized four-dimensional tensor; The standardized preprocessing step comprises: Spatial registration of images between different modalities or phases to align anatomical structures; Resampling the images to isotropic voxel resolution; Intensity normalization of the images to reduce the gray scale distribution difference introduced by different devices or scanning protocols; and Uniform the number of time phases of all image sequences to a preset length by interpolation or downsampling.
[0032] Specifically, in this stage, the original image data of different sources and formats is converted into unified and standardized model input, as follows: 1. Data collection and format unification: this method supports processing standard medical image formats such as DICOM and NIfTI. The system first parses the image metadata, extracts information such as examination time, device parameters, and performs de-identification processing to protect patient privacy. All data from different sources are converted into a unified internal data structure.
[0033] 2. Spatial registration and resampling: To eliminate spatial displacement between different modalities and different phases (e.g. caused by patient breathing, motion), the method adopts a two-stage registration strategy of “rigid + non-rigid”. First, a rigid registration based on mutual information maximization is used to align the global coordinate system. Then, a non-rigid registration model based on B-Spline is used to fine-tune the local deformation. After registration, all image volume data are resampled to isotropic voxels (e.g. 1.0 mm x 1.0 mm x 1.0 mm) by trilinear or higher order interpolation methods (e.g. B-Spline interpolation), and a fixed size (e.g. 128x128x128 voxels) region containing the target organ is cropped as input.
[0034] 3. Intensity normalization: To solve the gray value difference caused by different devices and protocols, the Z-score normalization method is used to subtract the mean value of each image volume data from its voxel value and divide it by its standard deviation, so that the gray value of all input images is distributed in the range of mean value 0 and variance 1. This operation preserves the relative intensity change (e.g. signal change of dynamic enhancement) while eliminating the systematic bias of absolute intensity.
[0035] 4. Time phase alignment: The acquisition timestamps in the metadata are used to sort the phases of the dynamic sequence. For sequences with different number of phases, linear interpolation or Gaussian process interpolation methods are used to unify the number of phases of all sequences to a preset fixed length T (e.g. T=10), ensuring the comparability of different cases in the time dimension.
[0036] 5. Data quality control: By calculating the signal-to-noise ratio (SNR), blurring degree and other indicators, low-quality phase frames or entire modalities are automatically screened and marked. These quality information can be used in the subsequent fusion stage.
[0037] After preprocessing, the original multi-modality dynamic images are converted into a set of standardized four-dimensional tensors .
[0038] Step S2: using a phase-aware feature encoder to extract multi-scale spatio-temporal features for each modality from the four-dimensional tensors; The step of extracting multi-scale spatio-temporal features by the phase-aware feature encoder comprises: dividing the volume data of each phase into three-dimensional voxel blocks, and embedding phase position encoding and spatial position encoding for each voxel block; constructing and fusing the difference features or ratio features of the voxel blocks between adjacent phases to enhance the representation of temporal dynamic changes; and processing the voxel block embedding using a stacked multi-head self-attention module, wherein the self-attention adopts a windowing strategy in the time dimension.
[0039] Specifically, this stage aims to efficiently extract features that contain spatial structure and temporal dynamics from the standardized four-dimensional tensor, as follows: 1. Voxel patch segmentation and spatio-temporal embedding: The volume data X^(m)_t (t=1,...,T) of each phase is divided into non-overlapping three-dimensional voxel patches (Patch) with a size of P×P×P. Each Patch is flattened and obtains initial tokens through a linear mapping layer. To introduce spatio-temporal position information, a learnable spatial position encoding and phase position encoding are added to each token.
[0040] 2. Temporal change enhancement: To explicitly capture dynamic changes, this method calculates the difference features and ratio features of the Patches at the same spatial position between adjacent phases t and t-1. After independent linear mapping, the two features are concatenated with the original tokens in the channel dimension. The difference features are sensitive to rapid rises and falls of signals, while the ratio features are not sensitive to absolute intensity. The two features complement each other and greatly enhance the model's ability to recognize perfusion patterns.
[0041] 3. Multi-head self-attention encoding: The enhanced token sequence is input into a Transformer-based encoder. The encoder consists of multiple stacked blocks, each containing a multi-head self-attention (MHSA) layer and a feed-forward network (FFN).
[0042] Temporal window attention: To reduce computational complexity and focus on relevant temporal dependencies, the attention calculation in the time dimension is limited to a local window centered on the current phase (e.g., [t-k, t+k]) when calculating attention. This captures key short-term dynamics while avoiding the massive computational overhead of 4D global attention.
[0043] Spatial sparse / window attention: Similarly, in the spatial dimension, windowed self-attention (W-MSA) and shifted windowed self-attention (SW-MSA) in Swin-Transformer or other sparse attention patterns can be used to achieve efficient long-range spatial information interaction.
[0044] Pyramid structure: The encoder adopts a pyramid structure to generate multi-scale feature maps with different spatial resolutions through progressive downsampling (e.g., Patch Merging) to capture both fine details and global structures.
[0045] Step S3: Through a modal confidence attention fusion mechanism, the multi-scale spatio-temporal features are adaptively weighted and fused according to the confidence estimated in real time for the features of each modality, generating fused features; The modal confidence attention fusion mechanism includes: For each modality, a quality assessment subnetwork is utilized to estimate a local confidence map reflecting local data quality and a global confidence score reflecting overall data quality from its spatio-temporal features; Based on the local confidence map and global confidence score, a fusion weight for each modality at each spatio-temporal location is calculated; and The features of each modality are weighted and summed according to the fusion weight to obtain the fusion features.
[0046] Referring to Figure 3 This stage aims to intelligently fuse features from multiple modalities to address the problem of uneven data quality in the real world.
[0047] Step S4: inputting the fusion features into one or more downstream task heads to generate diagnostic results; The downstream task head includes at least one of a segmentation head, a classification head, a detection head, or a regression head; The diagnostic results include a segmentation mask of a lesion, a classification probability of a disease, a detection bounding box of a lesion, or a voxel-level physiological parameter map.
[0048] In addition, a step of post-processing the diagnostic results is performed, and the post-processing includes: Temperature scaling calibration is performed on the classification probability; Temporal smoothing is performed on the segmentation results or parameter maps of dynamic sequences; and The diagnostic results and / or the modality confidence map are packaged and output in a format conforming to the DICOM standard.
[0049] Wherein, the parameters of the feature encoder and the fusion mechanism are determined through a training process comprising a self-supervised optimization task, and the self-supervised optimization task includes: A three-dimensional mask reconstruction task based on partial area occlusion of input images or features and prediction of occluded area content; and A phase sequence contrast learning task based on distinguishing correct phase sequence and incorrect phase sequence.
[0050] The training process further includes stability and generalization constraints, and the constraints are selected from at least one of the following: Deformation equivariance constraint, which requires the model to maintain equivariance in the feature space for geometric deformation applied to the input image; Confidence consistency constraint, which requires the model to maintain consistency in the estimated modality confidence for different views or perturbed versions of the same image.
[0051] Referring to Figure 2 In the model training stage, especially in the pre-training stage, the method learns a powerful feature representation without human annotation by using the following self-supervised tasks and constraints.
[0052] Three-dimensional mask reconstruction task: randomly mask a high proportion (e.g., 75%) of three-dimensional voxel blocks of the input image. The model can only see the unmasked part, and its task is to reconstruct the original voxel values of the masked area. This task forces the model to learn strong spatial context relationships and long-range dependencies, and the loss function is the mean square error (MSE) or L1 loss of the reconstructed area and the original area.
[0053] Phase sequence contrast learning task: this task aims to make the model understand the inherent temporal sequence of dynamic sequences. The features of adjacent phases (e.g., t and t+1) at the same spatial position are considered as a positive sample pair. The out-of-order phase pairs (e.g., t and t+3) or phase feature pairs from different cases are constructed as negative samples. The InfoNCE loss function is used, and the goal is to bring the positive sample pairs closer in feature space while pushing the negative sample pairs apart.
[0054] Deformation equivariance constraint: a random, differentiable geometric deformation T (including affine and non-rigid deformation) is applied to the input image. This constraint requires that the features obtained by first deforming the image and then encoding should be consistent with the result of first encoding and then deforming the feature map accordingly. This makes the model's learned features insensitive to small changes in position and shape, enhancing the robustness to different patient positions and organ movements.
[0055] Confidence consistency constraint: two different "views" (e.g., applying different noise or contrast changes) are generated for the same sample. This constraint requires that the confidence maps evaluated by the model for the same modality of these two views should remain highly consistent. This can stabilize the behavior of the MCA module and prevent it from producing dramatic weight changes due to small input disturbances.
[0056] The final self-supervised total loss function is the weighted sum of the above items: The overall loss function is:
[0057] where λ1, λ2, λ3, λ4 are adjustable weights, and the specific values can be determined by validation set tuning or automatic balancing methods. Through joint optimization, the network simultaneously obtains spatial structure modeling capability, temporal phase modeling capability, geometric disturbance stability, and cross-modality consistency on unlabeled data, so that it can still achieve high-level diagnostic performance under limited labeling; see the specific calculation process and formulas in Example 2 for details.
[0058] The model is first pre-trained on a large amount of unlabeled data using L total , and then fine-tuned on a small amount of labeled data combined with task-specific supervised loss (such as cross-entropy loss or Dice loss).
[0059] In summary, the scheme proposed by the present application can improve the dynamic perfusion analysis and disease change capturing capability. By introducing the MCA mechanism, the robustness of multi-modal image fusion is enhanced, and the interference of low-quality or abnormal modalities on the results is reduced. Through 3D mask prediction and temporal contrast learning, the dependence on artificial annotation is reduced, and the generalization performance under cross-center and cross-device conditions is improved. For rare diseases and long-tail clinical cases, the feature expression and contrast discrimination ability are optimized, and the sensitivity and specificity of diagnosis are improved.
[0060] Embodiment 2 By acquiring multi-modal three-dimensional dynamic medical image data, one or more modalities of three-dimensional dynamic medical image sequences are acquired, and the sequences contain multiple phases arranged in time sequence. Standardized preprocessing is performed on the image sequences to generate a standardized four-dimensional tensor. Then, through a self-supervised optimization task (running through feature modeling and fusion), a phase-aware feature encoder is used to extract multi-scale spatio-temporal features for each modality from the four-dimensional tensor. Through a modality confidence attention fusion mechanism, the confidence of each modality is estimated in real time based on the features, and the multi-scale spatio-temporal features are adaptively weighted and fused to generate fusion features. Then, phase difference and ratio feature enhancement (which can be embedded in the encoding or fusion stage) is added, and the fusion features are input into one or more downstream task heads to generate diagnostic results. Diagnosis prediction and positioning, and output and clinical report generation. The parameters of the feature encoder and the fusion mechanism are determined through a training process that includes a self-supervised optimization task, which includes a three-dimensional mask reconstruction task based on partially occluding regions of the input image or features and predicting the contents of the occluded regions, and a phase sequence contrast learning task based on distinguishing correct phase order and incorrect phase order.
[0061] According to the method of the first aspect, a specific embodiment is given as follows: I. Acquire one or more modalities of three-dimensional dynamic medical image sequences, which contain multiple phases arranged in time sequence. Standardized preprocessing is performed on the image sequences to generate a standardized four-dimensional tensor. First, in the aspect of data collection and format unification, the application supports multiple common medical image storage formats, mainly including DICOM standard format and NIfTI (.nii / .nii.gz) format based on scientific research. Clinical centers usually output DICOM file sequences, each sequence corresponding to one scan, containing complete metadata tags. The system of the application will call the DICOM parsing module to extract patient identification, examination time, device model, scan parameters and other key information in the import stage, and perform de-identification processing before storage, to ensure compliance with clinical ethics and privacy protection requirements. In the NIfTI format commonly used in scientific research data sources, the application also provides an interface for reading and automatically completing the missing metadata, so as to unify the subsequent processing process. Through this link, it is ensured that multi-modal dynamic images from different data sources can be received and analyzed under a unified input interface.
[0062] In the spatial registration and resampling stage, the application first processes the spatial differences across modalities and phases. The registration process adopts a two-stage strategy of "rigid registration + non-rigid registration", rigid registration is used for unifying the overall coordinate system, including translation, rotation and scale adjustment; non-rigid registration is used to eliminate local differences caused by organ movement, respiratory movement or patient position change. For rigid registration, the application adopts a registration method based on mutual information maximization, the calculation formula is:
[0063] Among them, A and B are the gray scale distributions of the two images, p ( a , b ) are joint probability distributions, p ( a ) and p ( b ) are marginal probability distributions. By maximizing mutual information, it is ensured that different modalities have the maximum correlation in statistical distribution, thereby realizing cross-modality alignment. Subsequently, in local non-rigid registration, the application uses a deformation model based on B-Spline, which deforms the image through a control point grid to obtain a continuous and reversible spatial mapping function, thereby eliminating local differences while ensuring the rationality of anatomical structures. After registration is completed, all images are mapped to a standard anatomical coordinate system.
[0064] To ensure the size consistency of network input and the comparability between modalities, the present application performs a resampling operation after registration. The three-linear interpolation is adopted to resample the volume data of different resolutions into isotropic voxels, and the voxel spacing is usually selected to be 1.0-1.5mm, which not only ensures that the spatial resolution is sufficient to capture small lesions, but also controls the data volume within the acceptable range of GPU memory. After resampling, the present application will perform cropping at the center position of the patient volume, retaining a cubic region containing the target organ and surrounding anatomical structures, and the cropping size is different according to different tasks, which is 112 3 to 160 3 between voxels. This cropping not only reduces the interference of redundant background, but also effectively reduces the consumption of computing resources.
[0065] In the intensity normalization and contrast adjustment link, the present application first considers the gray value difference caused by different devices and different scanning protocols. In order to ensure the comparability of cross-center images in intensity distribution, the present application adopts a two-step intensity standardization strategy. The first step is histogram matching, which matches the gray histogram of the input image with the histogram of the reference template image, so that different images remain consistent in global intensity distribution. The second step is Z-score-based normalization, whose formula is:
[0066] wherein, is the original voxel value, and are the mean and standard deviation of the voxel of the image, respectively, is the normalized value. Through this processing, the gray values of all input images are standardized to a distribution with a mean of 0 and a variance of 1, greatly reducing the differences in brightness and contrast between different center data. The above-mentioned image sequence is standardized and preprocessed in the above-mentioned manner to generate a standardized four-dimensional tensor. At the same time, for dynamic contrast-enhanced images, due to the perfusion of contrast agents, the signal intensity difference between different phases is significant, and the present application will retain the relative intensity change during normalization to avoid eliminating clinically meaningful dynamic information.
[0067] Temporal phase sorting and alignment are crucial steps in processing dynamically enhanced images in this invention. Different scanning devices may have different time intervals when acquiring phases, and some phases may even be missing or duplicated. If the phases are not sorted and aligned, subsequent temporal modeling will suffer from serious deviations. This invention first uses the trigger time or acquisition timestamp in the DICOM tag to sort the phases in chronological order; when complete tags are lacking, the phase order is inferred through cross-correlation analysis between images. For the sorted phase sequence, this invention uses linear interpolation or a Gaussian process-based temporal interpolation method to unify the number of phases for all patients to a fixed length T. For example, if the target length is set to 10, regardless of whether the original phases are 8 or 12, they are all unified to 10 phases through interpolation or downsampling. This phase alignment operation ensures the comparability of different patients in the temporal dimension.
[0068] In the data quality control stage, this invention adds quality screening for phase frames and modes to ensure the reliability of input data. Frames below a threshold are marked by calculating the signal-to-noise ratio (SNR), ambiguity index, and motion artifact score. When the SNR of a certain phase falls below a set threshold... θ snr Or the ambiguity index exceeds the threshold θ blur If the phase is deemed too low, it will be discarded or replaced by interpolation using a neighboring phase. If the overall quality of the entire mode is too low, its confidence weight will be automatically reduced in subsequent fusion stages to prevent low-quality modes from dominating diagnostic results. All quality scoring results will be logged for traceability and quality management.
[0069] 2. Using a phase-aware feature encoder, multi-scale spatiotemporal features are extracted from the four-dimensional tensor for each mode; Specifically, based on the standardized four-dimensional tensor obtained above, we will present an encoding method that can simultaneously express three-dimensional anatomical structure and temporal phase information. Let the first... m The dynamic volume data of each mode in a unified coordinate system are as follows: ,in T This represents the number of phases aligned in time sequence. To preserve the fine-grained spatial details of the volume data and the dynamic relationships between adjacent phases within acceptable memory and latency, this invention employs an encoding path of "voxel block segmentation—phase position embedding—phase difference / ratio enhancement—multi-head self-attention," outputting multi-scale, time-sensitive feature representations to provide stable input for subsequent modality confidence fusion and task head prediction.
[0070] During the voxel block segmentation stage, the volume data is divided into isotropic voxels of size [size missing]. P × P × PA 3D patch, indexed by phase in the time dimension. t ∈{1,…, T} Frame-by-frame sampling. For the first... t Phase 1 k Let there be *n* patches, and denote their flattened vectors as *p*. Through linear mapping Obtain the initial embedding To explicitly characterize the time sequence, a length of [length missing] is introduced at each phase. d Phase position vector p t and spatial location encoding s (k) The initial tokens formed by adding them together to create a spacetime mixture are mathematically represented as follows:
[0071] in p t Learnable parameters or deterministic encoding based on trigonometric functions can be used; s (k) Press Patch in ( H , W , D The computational calculation of grid coordinates in the space. This additive injection enables the encoder to have differentiable recognition capability for the alignment relationship of "same spatial location across phase" without changing the dimension and computational cost.
[0072] Relying solely on position embedding is insufficient to characterize the rapid dynamic changes during the enhancement process. Therefore, this invention simultaneously introduces the difference and ratio channels of adjacent phases in the feature construction to amplify the separability of temporal variations. A phase difference is defined for adjacent phases within the same spatial patch. Phase ratio Where ε>0 is a numerical stability constant, used to avoid instability caused by the denominator approaching zero. The difference term is more sensitive to the intensity transitions during the infusion rise and backwash stages, while the ratio term maintains scale invariance when the intensity calibrations of different equipment are inconsistent. Both are mapped by a linear mapping Δ isomorphic to the main channel. E , R After embedding, splicing is performed in the channel dimension, and then combined with... p t , s (k) Together, they form the initial token for time-series enhancement:
[0073] Through this construction, the network obtains an explicit prior on the "trend of adjacent phase intensity" before entering the attention calculation, making it easier for the downstream attention layer to capture time-related structural changes.
[0074] The backbone encoding adopts a structure of three-dimensional multi-head self-attention- feedforward network (MHSA-FFN) stack, and is assisted by residual and layer normalization to maintain training stability. Let the input of the i-th layer be l Then the single-head attention is calculated according to
[0075] The actual implementation adopts multi-head parallelism and concatenation in the channel dimension. To establish a dependence that is neither too local nor too globally diluted in the time dimension, the application implements a "short-medium" phase window strategy on the key-value selection of attention, that is, for a query from a token in phase t , only allows it to interact with the key-value in the range of {t 2,…,t+2}. When the time delay is relatively small or the disease exists for a longer time, the window can be expanded to {t T 3,…,t+3}, without using a complete full connection to the full phase, thereby preserving the sensitivity to the key time span at the cost of linear approximation. In the spatial dimension, the attention adopts a block sparse strategy or windowed self-attention within a three-dimensional neighborhood to reduce complexity, and realizes layer-by-layer aggregation of long-range spatial information through a shift mechanism across windows.
[0076] In order to take into account different scales of anatomical structures, such as small lesion boundaries and organ-level contours, the application stacks multiple levels of encoding in a pyramid manner, and updates the resolution and channel number at each level by using down-sampling convolution or patch merging operation with a stride of 2, while maintaining the phase window attention unchanged at each scale to ensure that the time sensitivity is not erased after down-sampling. The output of the i-th level encoding is denoted as l where (x H l , W l , D l ) is the spatial resolution, C l and c is the channel number. In order to facilitate subsequent cross-modal fusion, the features are kept consistent with the input in the time dimension, and the statistics are unified by layer normalization at each scale to avoid channel drift caused by modal differences.
[0077] In the feedforward network part after the multi-head self-attention, two layers of channel-by-channel multilayer perceptron (MLP) and GELU activation are adopted, the first layer expands the dimension to , and the second layer projects back to C l , DropPath is inserted to improve the generalization ability. Considering the coexistence of sparse boundaries and uniform regions in three-dimensional body data, the application adds a gating coefficient based on phase difference energy at the input end of the FFN to adaptively improve the nonlinear expression ability of "dynamic significant regions" and suppress the overfitting tendency of the background; σ is Sigmoid, β is a learnable temperature. The gated forward output is added to the attention output through the residual path to form a hierarchical stable spatiotemporal coding result.
[0078] To further stabilize the semantic consistency of "the same spatial position across phases", the application introduces a lightweight phase alignment constraint in the middle layer of the encoder. Let the phase sequence feature of the l layer at the spatial position u be , then smooth it in the time dimension through one-dimensional convolution φ and minimize the residual norm of adjacent phases, which is in the form:
[0079] This regularization term is not part of the main loss during training, but is used as a stabilizing term in the encoder during backpropagation, avoiding unnecessary jitter of features over time in the presence of noisy phases or slight misalignments, thereby improving the discriminability of subsequent phase sequence contrastive learning.
[0080] To control memory usage and keep the inference time controllable end-to-end, the encoder uses "interleaved arrangement" by phase block in implementation. Specifically, it maintains the continuity of phase indexes within the same batch to make attention queries more concentrated on hitting the key-value of adjacent phases, thereby reducing cache misses caused by sparse access; at the same time, it reuses the phase position vector p t only when the spatial position encoding s (k) changes. This implementation detail does not affect the expression power, but allows the T value to be larger under the same graphics memory conditions, meeting the clinical time series length requirements of dynamic enhancement sequences.
[0081] After encoding through several levels, a multi-scale feature set m oriented to a single modality is obtained . In scale alignment, the application does not pool the time dimension to avoid weakening the time information in pyramid fusion; if the task needs to be aligned with the voxel-level output, it can be restored to high resolution in the spatial dimension through tri-linear upsampling, while keeping TThe encoder remains unchanged. To enable subsequent modality confidence attention to be integrated into the quality assessment subnetwork, it also outputs a global temporal summary vector at each scale. ,in ψ For lightweight one-dimensional temporal convolution, Pool u The summaries are spatially global averages. In the fusion phase, these summaries are used to estimate modal-level confidence and guide the weighting of cross-modal attention.
[0082] Based on the output of phase-aware 3D encoding, a confidence-driven fusion mechanism for multimodal dynamic images is presented next. Let the... m The mode in the th ... l The coding features of each scale are Spatial location u In other words, the time phase is... t This paper addresses common issues in clinical data, such as cross-center and cross-device discrepancies, individual phase motion artifacts, and enhancement failures, which can lead to significant inconsistencies in the usability of different modalities across different spatial and temporal locations. Indiscriminate averaging or splicing will inevitably amplify the noise of low-quality modalities and weaken discriminative power. To address this, the present invention proposes Modal Confidence Attention Fusion (MCA). Its core idea is that a quality assessment subnetwork generates learnable confidence priors at both local and global levels, and injects these priors in a differentiable manner into cross-modal fusion. This allows the fusion weights to adaptively change with spatial and temporal location, and in missing or distorted scenarios, the weights are reduced or even masked to ensure the stability and interpretability of downstream tasks.
[0083] In constructing the confidence scores, MCA also uses local confidence plots. With global confidence Local confidence is determined by the quality assessment subnetwork. Q It is learned from the encoded features and their temporal derivatives. Specifically, it involves learning the features at each scale. Three complementary quality channels are applied: one is a signal-to-noise ratio approximation based on local feature statistics, using a three-dimensional sliding window to estimate the mean. μ with standard deviation σ ,form Second, a time-series stability measure based on phase difference and ratio is defined. To suppress severe discontinuities; and thirdly, to characterize the discernibility of boundaries and details based on texture integrity using three-dimensional gradient energy. These metrics are normalized and concatenated with the original features along the channel dimension, then input into a lightweight network consisting of two 3 × 3 × 3 depthwise separable convolutions and one 1 × 1 × 1 pointwise convolution. Q Output quality logarithm , and the local confidence is obtained by Sigmoid
[0084] The global confidence is fused by the encoder from the temporal summary vectors at each scale The present application adopts a gated weighted average pooling to form a cross-scale summary The weight is generated by a learnable soft selection, followed by an affine transformation and Sigmoid Intuitively, reflects the reliability prior of the modality in the whole sequence, while captures the local spatial-temporal availability, both of which will jointly determine the distribution of the fusion weight.
[0085] To inject confidence into the cross-modal information synthesis, the present application adopts a two-stage structure of "first intra-modal aggregation, then cross-modal fusion", which not only avoids exponentially expanding the key-value set in attention calculation, but also facilitates the implementation of explicit probability interpretation on the fusion weight.
[0086] In the first stage, three-dimensional-temporal self-attention is independently performed for each modality to obtain context-enhanced intra-modal features To reflect the constraint effect of local confidence within the modality, the present application does not directly multiply the confidence on the attention soft-max, but re-labels the intensity of the key-value pair, specifically:
[0087]
[0088] wherein represents a position-by-position and channel-by-channel Hadamard multiplication, avoiding excessive inhibition of gradients in the early training stage; the query branch Q is taken from the unweighted intra-modal projection (l) to prevent the formation of self-reinforcing closed-loop bias. The resulting has completed "soft suppression" of low-quality areas within the modality, laying the foundation for cross-modal fusion.
[0089] III. Through the modality confidence attention fusion mechanism, the confidence estimated in real time according to the features of each modality is used to adaptively weight and fuse the multi-scale spatio-temporal features, generating a fused feature; In the second stage, cross-modal fusion generates a set of normalized weights at each spatial-temporal position and performs weighted summation on the intra-modal features. To encode both global and local priors, the present application defines the logarithmic weight before fusion as Balance between two types of priors; introduce missing mask Indicate whether a modality exists or not. To avoid numerical instability in extreme cases where all modalities are judged as low quality, we use Softmax normalization with a smoothing term to get the final fusion weight
[0090] where is a tunable smoothing constant. This weight has a clear probabilistic meaning, when a modality is missing or of extremely low quality, its weight will naturally decay to close to zero. The final cross-modality fusion output is computed as
[0091] is computed and passed through a 1 × 1 × 1 linear rectification in the channel dimension to align the residual difference in the distribution of different modality channels. To further enhance the sensitivity of the fusion to the discrimination boundary, we apply a light temporal smoothing convolution with kernel length of 3—5 to suppress the fusion weight jittering caused by single-phase incidental noise; this smoothing is learnable in backpropagation, which helps to adaptively match the time constant of different diseases. m Light temporal smoothing convolution is applied to the beforehand to suppress the fusion weight jittering caused by single-phase incidental noise; this smoothing is learnable in backpropagation, which helps to adaptively match the time constant of different diseases.
[0092] In addition to cross-modality feature weighting, MCA also strengthens the key matching of reliable modalities by adding confidence bias to the attention log-likelihood. Specifically, a position-dependent bias term is added to the scoring matrix of attention within each modality to form
[0093] where κ >0 is a scaling factor. This term is equivalent to increasing the matching probability of high-confidence areas in the log space, which is particularly suitable for contrast-sensitive small perfusion change areas, allowing them to collect more information in the intra-modality context collection stage, and complementing the cross-modality probability weight m .
[0094] To avoid the fusion weight from excessively concentrating on a single modality at the early stage of training and falling into “collapse”, we introduce temperature regularization to the entropy of m in the total loss with a very small weight, i.e., minimizing the negative of , so as to maintain reasonable diversity without sacrificing interpretability. At the same time, the parameters of MCA and the quality evaluation subnetwork Q end-to-end receive joint gradients from task heads such as segmentation, classification, and regression as well as self-supervised objectives; when there are typical artifact frames or enhancement failure cases in the training samples, The stable global inhibition strategy, π m The weight of the corresponding region is reduced at the local level. It should be emphasized that the present application does not rely on external quality labels, but relies on the task loss and self-supervised loss to jointly shape the confidence output through the above-mentioned interpretable heuristic metric as an input channel, so that it maintains the "task relevance" consistent with the downstream diagnostic target.
[0095] The fusion feature is input into one or more downstream task heads to generate a diagnostic result. In terms of engineering implementation, the additional computational burden of MCA mainly comes from Q Several layers of lightweight convolution and two position-wise affine transformations, the time complexity is linearly related to the spatial resolution, and the memory usage is negligible compared to the main encoder. To ensure multi-scale consistency, the present application independently calculates and , and the is interpolated to high resolution at the same time to maintain the spatial alignment of the fusion weight and the feature; when the task head needs voxel-level output, The pyramid is up-sampled from top to bottom and the residual is added to the corresponding scale of the intra-modal feature, so that it can benefit from the enhancement of reliable modal and the fidelity of local details at high resolution. For extreme cases where only a single modality is available or all modalities are judged to be low quality, MCA will degenerate into the output of intra-modal attention or conservative fusion with smoothed minimum weight distribution, and through the system's quality alarm interface, it will prompt the outside to perform manual review or reacquisition at the clinical workstation.
[0096] The parameters of the feature encoder and the fusion mechanism are determined through a training process including a self-supervised optimization task, and the self-supervised optimization task includes: A three-dimensional mask reconstruction task based on partial region occlusion of input images or features and prediction of occluded region content; and A phase sequence contrast learning task based on distinguishing correct phase sequence and incorrect phase sequence.
[0097] After completing the modal confidence attention fusion, the system obtains a fused feature representation that has been aligned and weighted in space, time, and multi-modal dimensions. Although these features can better adapt to downstream segmentation, classification, or regression tasks, if completely relying on manually annotated samples for training, there are still two serious limitations. One is that the annotation of medical image data is extremely costly, and the annotation standards of different centers differ, resulting in insufficient cross-center migration performance. The second is that some tasks are inherently difficult to obtain fine-grained voxel-level annotations, such as subtle perfusion differences during dynamic enhancement or the ambiguous boundaries of early lesions, which are difficult to capture completely in manual annotation. To overcome the above problems, the present application introduces a self-supervised optimization task in the training phase of the fused features, uses unannotated or weakly annotated dynamic image data to construct learning signals, and drives the network to be more robust in time modeling and feature discrimination. The self-supervised module mainly includes a three-dimensional mask reconstruction task and a phase sequence contrast learning task, both of which are optimized on the same encoder and fusioner, and are supplemented by regularization constraints to form a joint loss function, so that the model can still obtain high-quality representation ability on a large amount of unlabeled data.
[0098] In the three-dimensional mask reconstruction task, the core idea is to randomly mask a part of the voxel blocks of the input tensor, so that the network predicts the content of the masked part only relying on the remaining visible area. This mechanism can force the model to learn the correlation between spatial neighborhoods and the dynamic consistency between time phases, thereby improving the robustness of the features to local missing and noise. In the specific implementation process, first, the voxel space where the fused features are located is randomly sampled, and a three-dimensional patch with a proportion of ρ is selected as the masking area. Let the input tensor be X , the masking area set be M , and the visible area be . The network encoder only processes V, and predicts through a lightweight decoder to reconstruct the masked patch. The loss function is defined as:
[0099] where M is the number of masked patches. This reconstruction error drives the network to capture the statistical rules between voxels at a global level through gradient backpropagation, thereby learning generalizable spatial features without relying on annotations. To avoid the model relying too much on low-frequency structures and ignoring detailed information, the present application adds a weight to the high-frequency area when calculating the loss, specifically multiplying a coefficient greater than 1 in the edge and texture dense area, so that the model pays more attention to the reconstruction of complex structure parts.
[0100] Mask reconstruction in the spatial domain is still insufficient to capture the temporal dependency of dynamic images, so the present application further proposes a phase order contrast learning task to explicitly enhance the model's sensitivity to time order. Different phases in dynamic images have natural order, for example, in the dynamic enhancement process, the signal rise in the early perfusion period and the signal fall in the late period should form a distinguishable time sequence track. If the model cannot correctly identify the logical relationship between adjacent phases, its generalization ability in disease discrimination will be greatly reduced. Therefore, the present application designs a phase order contrast learning mechanism based on InfoNCE loss. Specifically, let the features of the same spatial position in adjacent phases be z t and z t+1 , they constitute a positive sample pair; negative samples are generated by randomly shuffling the phase order or sampling across patients . The loss function is:
[0101] where sim is the cosine similarity function, τ is the temperature coefficient. By minimizing the loss, the model is forced to learn to make the positive sample pair close in the feature space, while the negative sample pair remains far away, thereby establishing an ordered representation of the time phase under the condition of no label. To avoid excessive dependence on short-range phase similarity during training, the present application introduces "mid-range disorder" in the negative sample sampling strategy, i.e. the disordered phases with a time span of 2 to 4 are used as additional negative samples, so that the network also has discriminability in longer time dependence.
[0102] In addition to these two core tasks, the present application also imposes regularization constraints on the training process to ensure the consistency and stability of the features. The equivariance constraint applies random rotation, affine or non-rigid deformation T( ) to the input data, and requires the network encoder to output features that satisfy equivariance, i.e. , where represents the encoder, represents the operation of performing the same deformation in the feature domain. Its loss function is:
[0103] This constraint ensures that the features learned by the network remain stable when facing common geometric disturbances, especially suitable for morphological shifts caused by device differences or patient position changes in cross-center data. At the same time, the consistency of confidence constraint compares the inputs of different views or enhancement methods and requires their distribution in modal confidence to remain consistent. Let the confidence of two views be and , then the loss is:
[0104] This ensures that the MCA module outputs stable weights under different conditions, thereby avoiding excessive fluctuations in confidence when there are slight differences in the input and improving generalization performance across devices and centers.
[0105] The final joint optimization objective consists of the above four parts: 3D mask reconstruction loss, phase order contrast learning loss, deformation isomorphic constraint loss, and confidence consistency constraint loss. The overall loss function is:
[0106] λ1, λ2, λ3, and λ4 are adjustable weights, and their specific values can be determined through validation set tuning or automatic balancing methods. Through joint optimization, the network simultaneously achieves spatial structure modeling ability, temporal phase modeling ability, geometric perturbation stability, and cross-modal consistency on unlabeled data, thus achieving a high level of diagnostic performance even with limited annotations.
[0107] During training, this invention typically performs self-supervised pre-training on a large-scale unlabeled dataset, relying solely on... L total Optimization was then performed. Fine-tuning was then carried out on a limited-label clinical task dataset, incorporating task-related supervision loss with... L total The self-supervised modules are combined and self-supervised constraints are retained with smaller weights, allowing for rapid adaptation to specific tasks without forgetting basic representations. During the inference phase, the self-supervised modules no longer incur additional computational burden because the mask reconstruction and contrastive learning branches are only enabled during training. In the final deployment, only the encoder, fusion unit, and task head are retained, ensuring inference efficiency and clinical feasibility.
[0108] Building upon the spatial and temporal feature representations obtained from the self-supervised optimization task, this section further introduces stability and generalization constraints. The aim is to ensure that the features remain robust and consistent across centers, devices, and protocols through deformation equivariance and confidence consistency mechanisms.
[0109] Deformation isomorphism constraints are achieved through a composite mapping Φ of the encoder-fusion unit. Let be the object, requiring that the spatial deformation suffered by the input end yields a consistent result in the feature domain through corresponding differentiable operations. T ( ) is from the deformable family Three-dimensional geometric transformation of mid-sample, It also includes small-angle rotations, anisotropic scale changes, affine transformations of translation, and non-rigid deformations controlled by B-Spline mesh parameterization. Let X The normalized four-dimensional input tensor (time dimension) obtained in the previous section TIf the loss is kept unchanged), then is the characteristic of "deformation before encoding", while is the result of implementing "encoding before deformation" in the feature domain with differentiable mesh resampling. We construct using trilinear interpolation and impose constraints on the outputs of each scale of the encoder to obtain the multi-scale isometry loss:
[0110] Unlike data augmentation in the input domain, this loss directly aligns the two forward paths in the feature space, making the network learn the priori that "what kind of geometric perturbation should not change the representation". In practice, we sample the affine parameters from a small range (e.g. rotation ±7 , scaling 0.9-1.1, translation no more than 5% of the body length), and the non-rigid control point displacement from zero-mean Gaussian sampling and Gaussian kernel smoothing to ensure smooth and reversible deformation. To avoid instability caused by feature domain folding, we add a soft constraint on the Jacobian determinant when generating the mesh to filter out sampling points that locally flip the voxel. Multi-scale summation makes the geometric boundaries of the shallow layer and the semantic structure of the deep layer simultaneously isometrically aligned, and thus remains robust in the cross-center scanning body position change and respiratory motion scenarios.
[0111] The consistency constraint of confidence is used to stabilize the weight distribution of the modality confidence attention (MCA) under different viewing angles and acquisition conditions, reducing the weight jitter caused by local noise and contrast changes. The local confidence and the global confidence are both learned end-to-end by the quality evaluation subnetwork. To construct a consistent signal, we generate two "viewing angle" versions, view1, view2, from the four-dimensional input of the same case during training. The two views share the same deformation T ( ) in geometry, and a light contrast / noise perturbation (e.g. linear contrast stretching and random pseudo noise) is applied in the intensity domain to simulate cross-device imaging differences. Let the local confidence calculated under the two views be and , then the space-time term of the consistency constraint is defined as:
[0112] where U l = H l W l D l. Considering the natural temporal continuity of adjacent phases in dynamic images, the present invention further introduces a temporal smoothing term within a single view, but avoids over-smoothing at the perfusion peak. To this end, the phase difference energy in the previous section is used as a gate to define the temporal constraint as:
[0113] where = exp( α mean∣Δ E (Δ X )∣( u , t )) to gate the strength of temporal smoothing by dynamic difference energy, and α > 0 as a temperature coefficient to make the low dynamic region more consistent while preserving the necessary changes in the high dynamic region. In addition, to avoid the fusion weight collapsing to a single mode, an entropy regularization (introduced in the previous section) is added to the aforementioned Softmax-based fusion weight π m ( u , t ) to form a comprehensive confidence consistency loss:
[0114] where is set by the validation set or automatically balanced. This combination stabilizes weight learning in three dimensions of space, time, and allocation entropy, maintaining preference for reliable modalities while avoiding the fragility brought by over-sharp weights.
[0115] To further resist the drift of inter-center contrast differences and noise statistics, this section adds a weak alignment strategy of feature matrix matching within the consistency framework, which is regarded as a special case of "view consistency" without introducing new independent losses. Specifically, the channel mean and variance are soft-constrained on the same scale and position of two-view features, and the Charbonnier distance is used to improve numerical stability, i.e., in the channel dimension of Φ (l) , μ (l) and σ (l) ,
[0116]
[0117] are added to the weight of L conf-st . This expression of "matrix alignment as consistency" avoids additional discriminators or domain adversarial modules, maintaining the simplicity of implementation and zero inference overhead.
[0118] In terms of training scheduling, the isometry and consistency constraints of the two categories are combined with theL MIM with L OCLL comprise the joint objective. Considering that imposing strong constraints too early can reduce the explorability of the self-supervised task, the present invention adopts a warm-up scheme: in the first few rounds of training, the L def with L conf , and then slowly increase its weight with learning rate decay, so that the network first acquires basic spatio-temporal representation, and then gradually tightens the stability and generalization boundary. The joint objective maintains the form of the previous section
[0119] where λ3, λ4 are linearly up-regulated from 0 or small values in the warm-up phase, and λ1, λ2 are slightly decreased in the middle and later stages to avoid competition between objectives. To suppress gradient resonance, the equivariant branch truncates the coordinate gradient sampled on the grid when computing , keeping numerical stability; the consistent branch avoids using the same random seed when constructing the two views, ensuring that the view difference has sufficient learning signal.
[0120] In the inference phase, no constraint term is explicitly calculated, and all costs have been internalized as parameters in training, so there is no increase in deployment latency. However, the "equivariance-consistency" prior learned in training will be reflected in two aspects: first, the encoder output is robust to small body position changes and non-rigid displacement, making the segmentation boundary and perfusion parameter regression insensitive to respiratory and motion artifacts; second, the weight distribution of MCA remains smooth and interpretable in the presence of contrast, noise, or partial modality absence, avoiding false dominance caused by incidental noise. To monitor these properties in engineering, the present invention measures the feature alignment error and weight jitter amplitude with paired deformation input and dual-view input in offline verification, taking the quantile statistics of as the "stability benchmark", and once it exceeds the threshold, it will backtrack the sample distribution or adjust the training weight, forming a closed-loop optimized production process.
[0121] The input in the inference phase is still a four-dimensional tensor after standardization, containing multi-phase body data from one or more modalities. When a modality is missing, the system will automatically adjust its weight to zero through the modality confidence mechanism, and the remaining modalities will continue to complete feature fusion and prediction. If the input volume size exceeds the memory limit, three-dimensional sliding window inference is used, with overlapping areas set between windows, and weighted averaging is used in the output stage to eliminate differences at block boundaries. Specifically, let the weight of the i th window on voxel u be w i ( u ), and the corresponding output isy i u
[0122] where the weight function takes value close to 1 at the center of the window and gradually decays at the boundary, thus ensuring smooth transition and no artifacts in the stitching result.
[0123] For segmentation tasks, the fused feature is passed through a dense prediction head to generate voxel-level probability maps at the highest resolution. Let the class set be {c}, the corresponding output probability is To overcome the confidence calibration difference between different centers, the invention introduces temperature scaling during inference, which adjusts the logits to get
[0124] where the temperature parameter τ is estimated on the validation set. In dynamic enhancement sequences, isolated phases may be incorrectly segmented due to noise, so a temporal smoothing strategy is further adopted. Specifically, the probability is integrated in the time dimension through an exponential moving average:
[0125] where α ∈[0.6,0.9] is used to control the smoothing strength. The smoothed probability map is thresholded to obtain the segmentation mask, supplemented by connected component filtering, volume lower bound constraint, and hole filling to ensure the result is coherent and consistent with anatomical rules. The final segmentation result can be single-phase output or full-time sequence output, and the peak phase or complete dynamic sequence can be selected for archiving according to the task requirements.
[0126] In detection and classification tasks, the fused feature is mapped to a case-level probability vector through global pooling and fully connected layers for disease classification or risk stratification. For three-dimensional detection tasks, candidate boxes are generated through dense anchor point or center point prediction, and consistent non-maximum suppression is performed in the time dimension to merge repeated candidates across phases. The suppression condition is determined by the joint score of three-dimensional IoU and cross-phase overlap rate. The posterior probability of classification and detection is also calibrated by temperature scaling to ensure that the threshold determination is consistent in different center usage scenarios. For cases that require graded risk output, the continuous probability value can be mapped to "low risk, medium risk, high risk" through a two-level threshold, and the threshold and model version number are recorded in the output report for subsequent quality tracing.
[0127] For parameter regression tasks such as perfusion parameter estimation, the model directly outputs continuous numerical maps at the voxel level To ensure physical plausibility, the last layer activation adopts Softplus or Sigmoid to ensure the output meets non-negative or within a set upper limit. If the task goal is a single parameter map rather than a time series, robust aggregation can be performed on the time dimension, such as taking the median , or weighted average , where the weights w ( t ) come from phase position or confidence to highlight phases with higher dynamic information. The final generated parameter map can be combined with the segmentation result to calculate the average, peak value, and time to peak (TTP) of the target region, and write them into the output file.
[0128] The explainability output in the inference stage mainly manifests in two aspects. The first is to obtain the probability variance by repeating forward calculation (such as enhanced or Monte Carlo Dropout during testing) to estimate the voxel-level prediction uncertainty. Let the probability of the jth forward be P (j) ( u , t ), then the variance is:
[0129] An uncertainty heat map can be formed to prompt the doctor to pay attention to the area where the prediction is not robust. The second is to directly output the modality confidence weight π m ( u , t ) and its spatiotemporal distribution obtained during the fusion process as visual evidence of the reliability of the information source, which is convenient for the doctor to understand the basis of the system's judgment.
[0130] All results are exported in DICOM compatible format to ensure seamless integration with existing in-hospital PACS / HIS systems. The segmentation results are exported as DICOM-SEG, accurately corresponding to the SeriesInstanceUID of the original image; the detection and classification results are recorded in the form of DICOM-SR, including class probability, risk level, model version, and confidence index; the parameter map output by the regression task is saved in the DICOM Parametric Map IOD format, accompanied by physical units and recommended window width and window level. All result files contain metadata, recording threshold settings, calibration parameters, and running environment information to support quality control and auditing.
[0131] In anomaly and degradation scenarios, the inference process possesses adaptive processing capabilities. When a modality is completely missing, its weight naturally resets to zero, and the system automatically continues diagnosis with the remaining modalities. When the overall input quality is low, the system still outputs prediction results, but adds a quality alarm flag to the report and retains the original probability map for manual review. If there is a memory overrun or some regions are not covered during inference, the result file will be marked as "partially completed" and include a retry flag and error log to prevent misuse of results. To improve batch inference efficiency, the system supports case-level asynchronous queue processing, with the full-phase sequence of each case pipelined for inference on the same GPU, ensuring timing consistency and resource utilization.
[0132] This invention proposes a self-supervised multimodal temporal fusion method for 3D dynamic medical image diagnosis. Starting with data preprocessing, it achieves robust feature extraction through phase-aware encoding and confidence-based attention mechanisms, and obtains cross-center and cross-modal generalization capabilities under joint optimization of self-supervised tasks and stability constraints. The inference stage utilizes a lightweight workflow to complete tasks such as segmentation, classification, and parameter regression. Results are output in a DICOM-compliant format, possessing interpretability and clinical applicability. The overall solution forms a complete closed loop from input to output, solving problems such as scarce annotations, modal differences, and insufficient generalization, providing a reliable implementation path for intelligent medical image diagnosis.
[0133] The second aspect of this invention discloses a three-dimensional dynamic medical image diagnostic system based on self-supervised multimodal temporal fusion. Figure 4 This is a structural diagram of a three-dimensional dynamic medical image diagnostic system based on self-supervised multimodal temporal fusion according to an embodiment of the present invention; as shown. Figure 4 As shown, the system 100 includes: A second aspect of this invention discloses a three-dimensional dynamic medical image diagnostic system based on self-supervised multimodal temporal fusion. The system employs the method described in either Embodiment 1 or Embodiment 2 of the first aspect. The system 100 includes: The data preprocessing module 101 is used to acquire and process one or more modal three-dimensional dynamic medical image sequences to generate standardized four-dimensional tensors. Feature encoding module 102 is used to extract multi-scale spatiotemporal features for each modality from the four-dimensional tensor using a phase-aware feature encoder; The multimodal fusion module 103 is used to adaptively weight and fuse the multi-scale spatiotemporal features based on the real-time estimated confidence of the features of each modality through a modal confidence attention fusion mechanism, thereby generating fused features. The diagnostic reasoning module 104 is used to input the fused features into one or more downstream task heads to generate diagnostic results; The parameters of the feature encoding module and the multi-modal fusion module are determined through a training process including a self-supervised optimization task. The self-supervised optimization task module 105 is configured to perform a three-dimensional mask reconstruction task of predicting the content of a blocked area by partially blocking an input image or feature, and a phase sequence contrast learning task based on distinguishing correct phase sequences and incorrect phase sequences.
[0134] The third aspect of the present application discloses an electronic device. The electronic device comprises a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the method for diagnosing a three-dimensional dynamic medical image based on self-supervised multi-modal time sequence fusion according to any one of the first aspect of the present application are implemented.
[0135] Figure 5 The structure of the electronic device according to the embodiment of the present application is shown in FIG. Figure 5 The electronic device comprises a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the electronic device is configured to provide computing and control capabilities. The memory of the electronic device comprises a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The communication interface of the electronic device is configured to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, operator network, near field communication (NFC) or other technologies. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen. The input device of the electronic device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the electronic device, or an external keyboard, touchpad or mouse, etc.
[0136] Those skilled in the art can understand that Figure 5 The structure shown in FIG.
[0137] The fourth aspect of the present application discloses a computer readable storage medium. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps of the method for diagnosing a three-dimensional dynamic medical image based on self-supervised multi-modal time sequence fusion according to any one of the first aspect of the present application are implemented.
[0138] Please note that the technical features of the above embodiments can be combined in any manner, and for the sake of brevity, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not contradict each other, they shall be considered within the scope of the present disclosure. The above embodiments only express several implementation manners of the present application, and the description is specific and detailed, but it shall not be understood as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these modifications and improvements shall be considered within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
[0139] The above are the preferred embodiments of the present application, and it should be noted that for those of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these modifications and improvements shall be considered within the protection scope of the present application.
Claims
1. A three-dimensional dynamic medical image diagnostic method based on self-supervised multimodal temporal fusion, characterized in that, Includes the following steps: Acquire a three-dimensional dynamic medical image sequence of one or more modalities, the sequence comprising multiple phases arranged in chronological order; The image sequence is preprocessed to standardize and generate a standardized four-dimensional tensor. Using a phase-aware feature encoder, multi-scale spatiotemporal features are extracted from the four-dimensional tensor for each modality; Through the modal confidence attention fusion mechanism, the multi-scale spatiotemporal features are adaptively weighted and fused based on the confidence estimated in real time according to the features of each modality, thereby generating fused features; Furthermore, the fused features are input into one or more downstream task heads to generate diagnostic results; The parameters of the feature encoder and the fusion mechanism are determined through a training process that includes a self-supervised optimization task, which includes at least one of the following: A 3D mask reconstruction task based on partially occluding the input image or features and predicting the content of the occluded area. A phase sequence comparison learning task based on distinguishing between correct and incorrect phase sequences.
2. The method according to claim 1, characterized in that, The standardized preprocessing steps include: Spatial registration is performed on images of different modalities or phases to align anatomical structures; Resample the image to isotropic voxel resolution; Image intensity normalization is performed to reduce grayscale distribution differences introduced by different devices or scanning protocols; and The temporal phase count of all image sequences is unified to a preset length by interpolation or downsampling.
3. The method according to claim 1, characterized in that, The steps of the phase-aware feature encoder to extract multi-scale spatiotemporal features include: The volume data of each phase is divided into three-dimensional voxel blocks, and phase position codes and spatial position codes are embedded in each voxel block. Constructing and fusing the differential or ratio features of the voxel blocks between adjacent phases to enhance the characterization of temporal dynamic changes; and The voxel block embedding is processed using stacked multi-head self-attention modules, wherein the self-attention employs a windowing strategy in the time dimension.
4. The method according to claim 1, characterized in that, The modality confidence attention fusion mechanism includes: For each modality, a local confidence map reflecting local data quality and a global confidence score reflecting overall data quality are estimated from its spatiotemporal features using a quality assessment subnetwork. Based on the local confidence map and global confidence score, the fusion weight of each modality at each spatiotemporal location is calculated; and The features of each modality are weighted and summed according to the fusion weights to obtain the fusion features.
5. The method according to claim 1, characterized in that, The training process also includes stability and generalization constraints, wherein the constraints are selected from at least one of the following: Deformation isovariability constraint requires the model to maintain isovariability in feature space for geometric deformations applied to the input image; The confidence consistency constraint requires that the model maintains consistent modal confidence for different viewpoints or perturbation versions of the same image.
6. The method according to claim 1, characterized in that, The downstream task head includes at least one of a segmentation head, a classification head, a detection head, or a regression head; The diagnostic results include lesion segmentation masks, disease classification probabilities, lesion detection bounding boxes, or voxel-level physiological parameter maps.
7. The method according to claim 6, characterized in that, It also includes a step of post-processing the diagnostic results, the post-processing including: Temperature scaling calibration is applied to the classification probabilities; Temporal smoothing of segmentation results or parametric maps of dynamic sequences; and The diagnostic results and / or the modal confidence plots are packaged and output in a format conforming to the DICOM standard.
8. A three-dimensional dynamic medical imaging diagnostic system, characterized in that, include: The data preprocessing module is used to acquire and process one or more modal three-dimensional dynamic medical image sequences to generate standardized four-dimensional tensors. The feature encoding module is used to extract multi-scale spatiotemporal features for each modality from the four-dimensional tensor using a phase-aware feature encoder. The multimodal fusion module is used to adaptively weight and fuse the multi-scale spatiotemporal features based on the real-time estimated confidence of the features of each modality through a modal confidence attention fusion mechanism, thereby generating fused features. The diagnostic reasoning module is used to input the fused features into one or more downstream task heads to generate diagnostic results. The parameters of the feature encoding module and the multimodal fusion module are determined through a training process that includes a self-supervised optimization task. The self-supervised optimization task module is used for 3D mask reconstruction tasks that partially occlude input images or features and predict the content of the occluded areas, as well as phase order contrast learning tasks that distinguish between correct and incorrect phase orders.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program. When the processor executes the computer program, it implements the steps of the three-dimensional dynamic medical image diagnosis method based on self-supervised multimodal temporal fusion as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the three-dimensional dynamic medical image diagnosis method based on self-supervised multimodal temporal fusion as described in any one of claims 1 to 7.