A Deep Learning-Based Multimodal Medical Imaging Teaching System and Method
Patent Information
- Application Number
- CN202611002407.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-07
- Publication Date
- 2026-09-01
AI Technical Summary
此类系统通常仅支持单一模态影像,无法有效整合计算机断层扫描、磁共振成像、超声等多种成像数据
本发明提供一种基于深度学习的多模态医学影像教学系统及方法。本发明通过空间对齐模块实现不同成像设备医学影像数据的体素级坐标统一,确保跨模态病灶位置精确对应,提升教学案例的解剖一致性。任务学习模型同步输出病灶分割掩码、病灶类别概率向量及像素级不确定性图,使人工智能的决策依据与认知局限可视化。可视化模块采用四窗格布局同步显示原始多模态影像、分割结果、梯度类激活映射热力图与不确定性分布图,并支持时间轴回溯推理过程,增强学生对AI工作机制的理解。参数更新模块在接收用户修正后仅更新解码器最后两层参数,既响应教学反馈又避免破坏主模型泛化能力。案例分级模块基于熵值、不确定性梯度总变差与病灶分割体积的比值及模态数量自动计算难度评分,实现教学资源的智能分级与适配。整体方案构建了从数据接入、智能分析、人机交互到案例组织的完整教学闭环,显著提升医学影像教学的临床真实性、可解释性与个性化水平。
Smart Images

Figure CN122676262A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the interdisciplinary field of medical education and artificial intelligence, and relates to a multimodal medical imaging teaching system and method based on deep learning. Background Technology
[0002] Current medical imaging teaching primarily relies on static image displays or playback functions based on traditional PACS systems. These systems typically support only a single modality of imaging and cannot effectively integrate multiple imaging data such as computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound. Even when some platforms attempt to introduce AI-assisted tools, the processing remains a black box, only providing the final lesion classification or segmentation results without a visual explanation of the reasoning behind it. Students struggle to understand how AI identifies lesions and cannot observe the model's uncertainties in areas with blurred boundaries or confused categories.
[0003] Furthermore, existing teaching systems generally lack interactive mechanisms, preventing users from correcting AI results or using feedback for model optimization. While some research has proposed multimodal image fusion methods, these often employ offline registration or simple image overlay, failing to achieve voxel-level spatial alignment. This leads to inconsistencies in lesion locations across modalities, impacting teaching accuracy. Other solutions attempt to introduce interpretability techniques such as gradient-based activation mapping, but these are not presented synchronously with segmentation results, uncertainty maps, and the original multimodal images, making it difficult to construct a complete cognitive chain. Simultaneously, the organization of teaching cases often relies on manual annotation of difficulty levels, which is inefficient and highly subjective, failing to dynamically assess complexity based on the model's output.
[0004] The aforementioned shortcomings make it difficult for existing technologies to support medical students' in-depth understanding and practical training in artificial intelligence-assisted diagnostic mechanisms. There is an urgent need for a new medical imaging teaching method that integrates multimodal alignment, interpretable reasoning, interactive correction, and intelligent grading. Summary of the Invention
[0005] To address the problems existing in the background technology, this invention proposes a multimodal medical imaging teaching system and method based on deep learning.
[0006] The first aspect of this application provides a deep learning-based multimodal medical imaging teaching system, comprising: The image access module is used to receive medical image data from different imaging devices; The spatial alignment module is used to map the medical image data to the same anatomical coordinate system through differentiable affine transformation and non-rigid cascade transformation, generating aligned multimodal images; The task learning model is used to process the aligned multimodal images and output lesion segmentation masks, lesion category probability vectors, and pixel-level uncertainty maps. The visualization module is used to simultaneously display the original multimodal image, the lesion segmentation mask, the heat map generated based on gradient class activation mapping, and the pixel-level uncertainty map, and supports users to correct the lesion segmentation results; The parameter update module is used to update only the parameters of the last two layers of the decoder in the task learning model based on the user-corrected segmentation results. The case classification module is used to calculate and classify the difficulty score of teaching cases based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
[0007] Optionally, the spatial alignment module uses one of the modal images as a reference modality, and the remaining modal images are aligned to the spatial coordinate system of the reference modality through a two-level transformation; the two-level transformation includes a first-level affine transformation and a second-level non-rigid transformation based on the displacement field.
[0008] Optionally, the task learning model adopts a shared encoder and multi-head decoder structure; the multi-head decoder includes a segmentation head, a classification head, and an uncertainty estimation head; the loss function of the uncertainty estimation head includes a product term of the pixel-level uncertainty map and the gradient magnitude of the segmentation mask.
[0009] Optionally, the visualization module uses a four-pane layout to simultaneously render the following: the original multimodal fusion image; a semi-transparent lesion segmentation mask; a gradient class activation mapping heatmap; a pixel-level uncertainty distribution map; and supports backtracking the evolution of intermediate features during the model inference process via a timeline slider.
[0010] Optionally, after receiving the segmentation mask corrected by the user, the parameter update module freezes the encoder parameters of the task learning model and performs gradient descent updates on the parameters of the last two layers of the decoder only based on the segmentation loss function.
[0011] Optionally, the formula used by the case grading module to calculate the difficulty score of the teaching case is: Difficulty score = α × entropy value + β × (total variation of uncertainty gradient / lesion segmentation volume) + γ × number of modes; Wherein, the entropy value is the information entropy of the lesion category probability vector; the total variation of the uncertainty gradient is the spatial gradient norm of the pixel-level uncertainty map; the lesion segmentation volume is the total number of voxels in the segmentation mask; the modality number is the number of modalities in the input image; and α, β, and γ are preset weight coefficients.
[0012] A second aspect of this application provides a deep learning-based multimodal medical image teaching method, including: Receives medical image data from different imaging devices; The medical image data is mapped to the same anatomical coordinate system through differentiable affine transformation and non-rigid cascade transformation to generate aligned multimodal images; The aligned multimodal image is processed using a task learning model to output a lesion segmentation mask, a lesion category probability vector, and a pixel-level uncertainty map. The system simultaneously displays the original multimodal image, the lesion segmentation mask, the heat map generated based on gradient class activation mapping, and the pixel-level uncertainty map, and receives user correction operations for the lesion segmentation results. Based on the user-corrected segmentation results, only the parameters of the last two layers of the decoder in the task learning model are updated; The difficulty score of the teaching case is calculated and graded based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
[0013] Optionally, the step of mapping medical image data to the same anatomical coordinate system includes: selecting one modal image as a reference modality; and sequentially performing affine transformation and displacement field-based non-rigid transformation on the remaining modal images to make their spatial coordinates consistent with the reference modality.
[0014] Optionally, the training loss function of the task learning model is composed of a weighted sum of segmentation loss, classification loss, and uncertainty loss; the uncertainty loss is the average of the product of pixel-level uncertainty value and the gradient magnitude of the segmentation mask at the corresponding position.
[0015] Optionally, the original multimodal fused image, semi-transparent lesion segmentation mask, gradient class activation map heatmap, and pixel-level uncertainty distribution map are presented through a four-pane interface; and a timeline control is provided to dynamically display the evolution sequence of feature maps during model inference.
[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention provides a deep learning-based multimodal medical imaging teaching system and method. The invention achieves voxel-level coordinate unification of medical image data from different imaging devices through a spatial alignment module, ensuring accurate correspondence of lesion locations across modalities and improving the anatomical consistency of teaching cases. The task learning model synchronously outputs lesion segmentation masks, lesion category probability vectors, and pixel-level uncertainty maps, visualizing the decision-making basis and cognitive limitations of artificial intelligence. The visualization module uses a four-pane layout to synchronously display the original multimodal images, segmentation results, gradient class activation mapping heatmaps, and uncertainty distribution maps, and supports time-axis backtracking inference processes, enhancing students' understanding of the AI's working mechanism. The parameter update module updates only the last two layers of the decoder parameters after receiving user corrections, responding to teaching feedback while avoiding disrupting the generalization ability of the main model. The case grading module automatically calculates difficulty scores based on entropy, the ratio of total variation of uncertainty gradients to lesion segmentation volume, and the number of modalities, achieving intelligent grading and adaptation of teaching resources. The overall solution constructs a complete teaching closed loop from data access, intelligent analysis, human-computer interaction to case organization, significantly improving the clinical authenticity, interpretability, and personalization of medical imaging teaching. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of a deep learning-based multimodal medical imaging teaching system according to an embodiment of the present invention; Figure 2 This is a flowchart of a deep learning-based multimodal medical image teaching method according to one embodiment of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] In one embodiment, such as Figure 1 As shown, a deep learning-based multimodal medical imaging teaching system is provided. This system corresponds one-to-one with the deep learning-based multimodal medical imaging teaching methods described in the following embodiments. The deep learning-based multimodal medical imaging teaching system includes: an image access module, a spatial alignment module, a task learning model, a visualization module, a parameter update module, and a case classification module. Detailed descriptions of each functional module are as follows: The system comprises the following modules: an image acquisition module for receiving medical image data from different imaging devices; a spatial alignment module for mapping the medical image data to the same anatomical coordinate system through differentiable affine transformations and non-rigid cascade transformations, generating aligned multimodal images; a task learning model for processing the aligned multimodal images, outputting lesion segmentation masks, lesion category probability vectors, and pixel-level uncertainty maps; a visualization module for simultaneously displaying the original multimodal images, the lesion segmentation masks, a heatmap generated based on gradient class activation mapping, and the pixel-level uncertainty maps, and supporting user correction of the lesion segmentation results; a parameter update module for updating only the parameters of the last two layers of the decoder in the task learning model based on the user-corrected segmentation results; and a case grading module for calculating and grading the difficulty of teaching cases based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
[0020] The image access module receives medical image data from various imaging devices, including computed tomography (CT), magnetic resonance imaging (MRI), ultrasound, and positron emission tomography (PET). The medical image data is transmitted in a digital medical imaging communication standard format, specifically image files conforming to the DICOM 3.0 protocol. The image access module acquires the image files via a network interface or local storage interface and parses the metadata information, including patient identification, examination time, device model, imaging parameters, spatial resolution, and voxel size. The image access module supports simultaneous access to image data from two or more modalities and assigns a unique session identifier to each set of image data to ensure that the correspondence between modalities is not confused during subsequent processing.
[0021] For example, a user uploads a data package containing chest computed tomography (CT) images, cardiac magnetic resonance imaging (MRI) images, and abdominal ultrasound images via a teaching terminal. The image access module automatically identifies the modality of each image and extracts its spatial dimensions: 512×512×120 voxels, 256×256×80 voxels, and 384×384×60 voxels, respectively. The image access module also performs preliminary format validation on the input images, removing files that do not conform to the DICOM standard or lack key metadata, thus ensuring the stability of subsequent alignment and analysis processes. Through the image access module, the system can seamlessly integrate heterogeneous medical images generated in real-world clinical environments, providing a data foundation for constructing multimodal joint teaching cases.
[0022] In this invention, the design of the image access module avoids the shortcomings of traditional teaching systems that only support a single modality or require manual image preprocessing, thereby improving the clinical authenticity of the teaching content and the ease of use of the system.
[0023] The spatial alignment module maps the medical image data to the same anatomical coordinate system through differentiable affine and non-rigid cascade transformations, generating aligned multimodal images. The spatial alignment module first selects one modality from the input multimodal images as a reference modality. The selection of the reference modality is based on clinical commonality or the principle of highest spatial resolution; for example, in a dataset containing computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound images, CT images are typically chosen as the reference modality. The remaining modal images undergo two levels of transformation processing to achieve spatial consistency with the reference modality. The first level transformation is a differentiable affine transformation, used to correct for differences in overall translation, rotation, scaling, and shearing. This differentiable affine transformation is parameterized by a fully connected neural network, outputting a 3×4 affine matrix that acts on the three-dimensional coordinate grid of the source image. The second level transformation is a displacement field-based non-rigid transformation used to compensate for local anatomical structural deformations. This non-rigid transformation predicts the three-dimensional displacement vector of each voxel through a convolutional neural network with an encoder-decoder structure, forming a dense displacement field, and resamples the source image using bilinear interpolation. The two-stage transformations are cascaded, meaning the output of the affine transformation serves as the input to the non-rigid transformation. The entire process is end-to-end differentiable and supports gradient backpropagation.
[0024] For example, the input MRI image has a size of 256×256×80 voxels, and the reference computed tomography image has a size of 512×512×120 voxels. The spatial alignment module first upsamples the MRI image to a physical space range similar to the reference modality, then coarsely aligns the chest cavity contour through affine transformation, and subsequently finely adjusts the position of the heart boundary using non-rigid transformation, so that the myocardial structure spatially overlaps in both modalities. All aligned modal images share the same voxel coordinate system, with a uniform voxel spacing of 1 mm × 1 mm × 1 mm. This design ensures anatomical consistency in cross-modal feature extraction during subsequent lesion analysis.
[0025] In this invention, through the design of the spatial alignment module, the system can automatically eliminate spatial deviations caused by different imaging devices due to acquisition angles, patient positions, or physiological movements. This ensures precise correspondence of the same anatomical region in multimodal images, providing a reliable foundation for demonstrating multimodal complementary information during teaching. It avoids the subjectivity and time-consuming nature of manual registration and overcomes the limitations of traditional rigid registration in handling organ deformation, significantly improving the accuracy and credibility of multimodal teaching cases.
[0026] The task-learning model processes the aligned multimodal images, outputting a lesion segmentation mask, a lesion category probability vector, and a pixel-level uncertainty map. The task-learning model employs a network architecture of a shared encoder and a multi-head decoder. The shared encoder consists of multiple stacked convolutional blocks, each containing a convolutional layer, a batch normalization layer, and an activation function, used to extract high-dimensional feature representations after multimodal fusion. The multi-head decoder includes three independent branches: a segmentation head, a classification head, and an uncertainty estimation head. The segmentation head progressively restores the spatial resolution through an upsampling path, ultimately outputting a binary lesion segmentation mask of the same size as the input image. The classification head performs global average pooling on the high-level feature map of the encoder, followed by a fully connected layer to generate a lesion category probability vector, the dimension of which is equal to the total number of preset disease categories. The uncertainty estimation head has a similar structure to the segmentation head, but outputs only one channel, generating a pixel-level uncertainty map with values ranging from 0 to 1, used to characterize the model's confidence level in the prediction results for each voxel. The model employs a weighted combined loss function during training, consisting of three parts: segmentation loss, classification loss, and uncertainty loss. The segmentation loss is expressed as 1 minus the Dice coefficient, the classification loss uses the cross-entropy function, and the uncertainty loss is defined as the average product of the pixel-level uncertainty value and the gradient magnitude of the corresponding segmentation mask, aiming to concentrate high-uncertainty regions at areas with blurred lesion boundaries. The total number of model parameters is kept below five million to adapt to general-purpose graphics processors in teaching environments.
[0027] For example, the input is aligned chest computed tomography (CT) and magnetic resonance imaging (MRI) images fused together, with a size of 512×512×100 voxels. The model output includes a 3D segmentation mask for the lung nodules, probability vectors for both benign and malignant nodules, and a pixel-level uncertainty map reflecting the degree of blurring at the nodule edges. The pixel-level uncertainty map is presented as a heatmap in the visualization interface, with red areas indicating that the model has significant doubt as to whether the area belongs to a lesion.
[0028] In this invention, through a task-based learning model, the system not only provides the location and type of lesions but also simultaneously reveals the model's own cognitive limitations. This design enables students to observe the reasoning basis and uncertainties of artificial intelligence in medical image analysis, thereby understanding the necessity of human-machine collaboration in clinical decision-making. This technology breaks through the limitations of traditional teaching systems that only display the final results, realizing a shift from black-box output to transparent reasoning, and enhancing the depth of understanding of AI-assisted diagnostic mechanisms in medical education.
[0029] It should be noted that the task learning model employs a shared encoder and multi-head decoder structure. The shared encoder receives aligned multimodal images as input and extracts multi-level spatial semantic features through progressively downsampled convolutional operations. This encoder consists of four encoding stages, each composed of two convolutional layers and one max-pooling layer. The output feature maps are then passed to the corresponding layers of the multi-head decoder. The multi-head decoder consists of three functionally independent sub-networks: a segmentation head, a classification head, and an uncertainty estimation head. The segmentation head uses a symmetrical upsampling path, progressively restoring spatial resolution through transposed convolutions and fusing skip connection features from the encoder to ultimately generate a lesion segmentation mask with the same size as the input image. The classification head starts from the deepest feature map of the encoder, first performing global average pooling to compress the spatial dimension, and then outputting a lesion category probability vector through two fully connected layers, with the number of elements equal to the preset total number of disease categories. The uncertainty estimation head has a similar structure to the segmentation head, also including upsampling paths and skip connections, but its output is a single-channel tensor. Each voxel value represents the uncertainty of the model's prediction of the lesion boundary at that location, with values limited to 0 to 1. During training, this uncertainty estimation head is constrained by a specific loss term, defined as the mean of the element-wise product of the pixel-level uncertainty map and the gradient magnitude of the lesion segmentation mask. The gradient magnitude of the lesion segmentation mask is calculated using the Sobel operator and is used to characterize the sharpness of the lesion boundary. For example, when the model processes lung nodule images, if the lesion boundary in a certain region is blurred or there is partial volume effect, the gradient magnitude of the segmentation mask at that location is low, and the uncertainty estimation head is guided to output higher uncertainty values in such regions. Conversely, in regions with clear boundaries, the gradient magnitude is high, and the uncertainty value is suppressed.
[0030] In this invention, through the mechanism of a task-based learning model, a pixel-level uncertainty map can accurately reflect the confidence distribution of the model's local predictions, thus making the internal judgment logic of artificial intelligence observable. Students can simultaneously observe lesion location, category judgment, and model confidence during the teaching process, thereby understanding why certain cases require physician review. This technology expands the model output from deterministic results to probabilistic cognition, significantly improving the interpretability and educational value of AI-assisted diagnosis in medical imaging teaching.
[0031] The visualization module synchronously displays the original multimodal images, the lesion segmentation mask, the heatmap generated based on gradient class activation mapping, and the pixel-level uncertainty map, and supports user correction of the lesion segmentation results. This module uses a four-pane layout to simultaneously present four types of information on the graphical user interface. The upper left pane displays the original multimodal images, using pseudo-color fusion to overlay grayscale images of different modalities into a single view, preserving the anatomical details of each modality. The upper right pane displays the lesion segmentation mask, with semi-transparent color blocks covering the original images, visually identifying the lesion regions determined by the model. The lower left pane displays the heatmap generated based on gradient class activation mapping, obtained by weighted summation of gradients from the encoder's high-level feature maps using backpropagation classification loss, highlighting the anatomical regions that contribute most to the classification decision. The lower right pane displays the pixel-level uncertainty map, using a red-yellow-blue color mapping to represent the level of uncertainty; red indicates that the model has significant doubt about whether the voxel belongs to a lesion. The four panes share the same spatial coordinate system and slice index. When a user scrolls the axial, coronal, or sagittal slice slider in any pane, the corresponding slices in the other panes are updated synchronously. The visualization module also integrates a timeline control, allowing users to drag the slider to trace the evolution sequence of intermediate features during the model inference process, such as the layer-by-layer changes from the initial convolution response to the final segmentation output. For correction operations, users can use a brush tool to add or erase marks on the segmentation mask; the system records the correction trajectory in real time and generates a new segmentation mask. All interactive operations are implemented through the WebGL rendering engine, ensuring smooth rotation, scaling, and slice browsing of the 3D image.
[0032] For example, a medical student observed a set of multimodal teaching cases on liver tumors and found that the model misidentified some blood vessels as lesions. After the student corrected this using the eraser tool in the upper right pane, the system immediately updated the displayed results and passed the corrected data to the subsequent processing module. This design allows students not only to passively receive AI output but also to actively participate in the diagnostic reasoning process. By synchronously comparing the original images, the model's judgment criteria, the confidence distribution, and the editable results, students can gain a deeper understanding of the working logic and limitations of artificial intelligence in medical image analysis.
[0033] In this invention, the visualization module effectively bridges the cognitive gap between theoretical teaching and clinical AI practice, and improves the training effect of human-machine collaborative decision-making ability in medical education.
[0034] It should be noted that the visualization module uses a four-pane layout to simultaneously render the following content: the original multimodal fusion image; a semi-transparent lesion segmentation mask; a gradient activation map heatmap; and a pixel-level uncertainty distribution map. Furthermore, it supports tracing the evolution of intermediate features during the model inference process using a timeline slider. The four panes are arranged in a two-dimensional matrix, with each pane occupying a quarter of the interface area to ensure clear visual contrast. The original multimodal fusion image is generated by overlaying aligned computed tomography (CT) images, MRI images, etc., using different pseudo-color channels, preserving the tissue contrast characteristics of each modality. The semi-transparent lesion segmentation mask is overlaid on the original image with a fixed transparency, allowing the underlying anatomical structures to remain discernible, facilitating the assessment of the degree of agreement between the segmentation boundaries and the actual lesions. The gradient activation map heatmap is generated by globally averaging the gradient of the classification task loss function with respect to the last layer of the encoder's feature map, then weighted and summed with the original feature map, and upsampled to the input resolution. Finally, it highlights the areas that play a crucial role in diagnostic decisions with warm colors. The pixel-level uncertainty distribution map uses a continuous color gradient to represent the prediction uncertainty of the model at each voxel location, transitioning from blue through yellow to red, corresponding to increasing uncertainty. The four panes share the same spatial coordinate system; when the user adjusts the slice position or viewpoint of any pane, the other panes automatically update synchronously to maintain consistent anatomical position. The timeline slider is located at the bottom of the interface, with its scale corresponding to key stages in the model inference process, including initial convolutional response, intermediate encoded features, skip connection fusion results, and final segmentation output. When the user drags this slider, the system dynamically replaces the content displayed in the four panes with the intermediate results of the corresponding inference stage. For example, when the slider moves to the third stage of the encoder, the gradient class activation map heatmap is updated to the attention distribution calculated based on the features of that stage.
[0035] For example, the teacher guides students to analyze a case of glioma in the brain, using a timeline slider to demonstrate how the model gradually focuses from early peripheral responses to the tumor core, while observing the increase in pixel-level uncertainty at the tumor infiltration margin. This design visualizes the abstract deep learning reasoning process. Students can intuitively understand how artificial intelligence extracts features layer by layer and forms diagnostic criteria.
[0036] In this invention, the visualization module transforms the black-box model into an observable, traceable, and interventionable teaching object, significantly enhancing the depth of understanding of artificial intelligence mechanisms and the sense of practical participation in medical imaging teaching.
[0037] The parameter update module updates only the parameters of the last two layers of the decoder in the task learning model based on the user-corrected segmentation results. This module initiates an incremental learning process upon receiving the corrected lesion segmentation mask submitted by the user through the visualization module. First, the system pairs the corrected segmentation mask with the original input image to construct a single-sample fine-tuning dataset. Then, it calculates the segmentation loss of the current model on that sample, expressed as 1 minus the Dice coefficient. During backpropagation, this module freezes all parameters of the first few layers of the shared encoder and decoder, performing gradient descent updates only on the weights and biases of the last two convolutional layers of the decoder. These two convolutional layers are located at the end of the upsampling path and directly determine the detailed shape of the final segmentation mask; therefore, fine-tuning them can effectively adjust boundary predictions without compromising high-level semantic feature extraction capabilities. The update process employs a small learning rate and a single-iteration strategy to avoid overfitting or parameter oscillations. The updated model parameters are saved as a dedicated version for the current teaching session, marked as a teaching revision, and used only for this course or the current user session, without affecting the global parameters of the system's main model.
[0038] For example, in a liver tumor case, students manually remove mis-segmented bile duct regions from the model. After submitting the corrected results, the parameter update module only adjusts the convolutional kernels of the last two layers of the decoder, enabling the model to more accurately distinguish between bile ducts and tumor tissue in subsequent slices, while maintaining its ability to recognize irrelevant anatomical structures such as the lungs or brain. This mechanism ensures immediate responsiveness to teaching feedback while maintaining the generalization stability of the main model. Through this design, the system achieves closed-loop teaching through human-machine collaboration, enabling the AI model to adapt to individualized teaching needs without relying on large-scale retraining.
[0039] In this invention, the parameter update module effectively solves the problem of static and fixed AI models in traditional teaching systems, which cannot respond to teacher-student interactions, thereby improving the flexibility and teaching adaptability of artificial intelligence tools in medical education.
[0040] It should be noted that after receiving the user-corrected segmentation mask, the parameter update module freezes the encoder parameters of the task learning model and only performs gradient descent updates on the parameters of the last two layers of the decoder based on the segmentation loss function. This process first verifies the spatial consistency between the user-submitted corrected segmentation mask and the original input image, ensuring that their voxel coordinates are aligned. Subsequently, the system constructs a single-sample training pair, containing the aligned multimodal image and its corresponding corrected segmentation mask. The segmentation loss function is in the form of 1 minus the Dice coefficient, used to measure the degree of overlap between the model's current output and the corrected mask. During the backpropagation phase, the gradient of this loss function with respect to all model parameters is calculated, but gradient updates are applied only to the parameters of the last two convolutional layers of the decoder. The parameters of all other layers, including those of the shared encoder, the front end of the decoder, and the classification head and uncertainty estimation head, are set to an untrainable state. The last two layers of the decoder typically consist of a 3×3 convolutional layer followed by a 1×1 convolutional layer, responsible for mapping the upsampled feature map to the final lesion probability distribution. The gradient descent update uses a fixed small learning rate and is limited to a single iteration to avoid overfitting to the single corrected sample. The updated parameters do not overwrite the main system model, but are stored as a temporary model version bound to the current teaching session.
[0041] For example, when explaining a case of lung nodules, the teacher discovers that the model misclassifies vascular branches as nodules. The teacher uses an interactive tool to erase the mis-segmented area and submits the corrected result. The parameter update module then freezes the weights of all convolutional blocks in the encoder, adjusting only the kernels of the last two layers of the decoder. This reduces the model's response intensity to similar vascular structures in subsequent slices of the same or adjacent slices. Because the high-level semantic feature extraction capability remains unaffected, the model's performance in recognizing other lesion types is unaffected. This mechanism achieves precise response to teaching feedback. Students can instantly see the impact of corrective actions on the AI output, thus understanding the dynamic process of human-machine collaborative diagnosis.
[0042] In this invention, the parameter update module avoids the risk of catastrophic amnesia caused by full model fine-tuning, while ensuring that the generalization ability learned by the main model on large-scale clinical data is not destroyed, so that the teaching system has both stability and interactive adaptability.
[0043] The case classification module calculates and classifies teaching cases based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities. The case classification module first obtains three core indicators from the task learning model. The first is the entropy value of the lesion category probability vector, calculated using the information entropy formula, reflecting the model's confidence in classifying lesions; a higher entropy value indicates greater uncertainty in the classification result. The second is the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume. The total gradient variation is obtained by summing the absolute values of the first-order differences calculated along the x, y, and z directions in three-dimensional space for the pixel-level uncertainty map. The lesion segmentation volume is the total number of voxels with a value of 1 in the segmentation mask. This ratio characterizes the relative relationship between the degree of lesion boundary ambiguity and the overall size. The third is the number of input modalities, i.e., the number of medical imaging modalities participating in the current case analysis. For example, when computed tomography, magnetic resonance imaging, and ultrasound images are included simultaneously, the number of input modalities is 3. The case grading module substitutes the three indicators mentioned above into a preset linear weighted formula to calculate the difficulty score of the teaching case. The weight coefficients in the formula are fixed constants, corresponding to the contribution ratios of classification uncertainty, boundary ambiguity, and multimodal complexity to the teaching difficulty. Based on the numerical range of the difficulty score, the system classifies the teaching cases into three levels: beginner, intermediate, and advanced.
[0044] For example, a case involving pulmonary ground-glass nodules exhibits a near-uniform distribution of lesion category probability vectors and high entropy. The pixel-level uncertainty map shows high gradient changes at the nodule edges, and the small lesion volume results in a large ratio of total gradient variation to volume. Furthermore, this case only contains a single modality of imagery. After comprehensive calculation, the difficulty score is medium, and it is categorized into the intermediate-level teaching case library. Another case involves multiple liver metastases, including both computed tomography (CT) and magnetic resonance imaging (MRI) modalities. The model clearly identifies some lesion categories, but the boundaries are highly ambiguous, and the gradient distribution of the pixel-level uncertainty map is complex, ultimately placing it in the advanced difficulty range. This grading mechanism is entirely based on the model's own output features, requiring no manual difficulty labeling. Teachers can automatically recommend teaching cases of appropriate difficulty levels based on students' grade level or training stage.
[0045] In this invention, the case grading module enables intelligent adaptation of teaching resources, allowing beginners to avoid overly complex, multimodal, and ambiguous cases, while senior students can challenge themselves with real clinical scenarios with high uncertainty. This technology enhances the personalization of medical imaging teaching and the scientific rigor of cognitive progression.
[0046] Specific limitations regarding the deep learning-based multimodal medical imaging teaching system can be found in the limitations of the deep learning-based multimodal medical imaging teaching method below, and will not be repeated here. Each module in the aforementioned deep learning-based multimodal medical imaging teaching system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0047] It should be noted that the formula used by the teaching case grading module to calculate the difficulty score of a teaching case is: Difficulty Score = α × Entropy Value + β × (Total Variation of Uncertainty Gradient / Lesion Segmentation Volume) + γ × Number of Modalities. Here, the entropy value is the information entropy of the lesion category probability vector, used to quantify the uncertainty of the model's judgment of the lesion category; this information entropy is calculated by summing the negative product of each category probability and its logarithm; the more uniform the probability distribution, the larger the entropy value. The total variation of uncertainty gradient is the spatial gradient norm of the pixel-level uncertainty map. It is calculated by taking the absolute difference of adjacent voxels along the x, y, and z directions in three-dimensional space for the pixel-level uncertainty map, and then summing the absolute differences in all directions, reflecting the drastic change of uncertainty in space. The lesion segmentation volume is the total number of voxels in the segmentation mask, i.e., the number of voxels predicted as lesions, measured in voxels. The number of modalities is the number of modalities in the input image; for example, when simultaneously inputting computed tomography (CT) images and magnetic resonance imaging (MRI) images, the number of modalities is 2. α, β, and γ are preset weighting coefficients used to adjust the relative importance of classification uncertainty, boundary ambiguity, and multimodal complexity in the overall difficulty. Their values are determined by the consensus of teaching experts and fixed in the program before the system is deployed, and are not dynamically adjusted with specific cases.
[0048] For example, a teaching case about a brain glioma includes three modalities: T1-weighted MRI, T2-weighted MRI, and enhanced MRI, for a total of three modalities. The model's output lesion category probability vector is nearly evenly distributed between high-grade gliomas and metastases, resulting in a high entropy value. The pixel-level uncertainty map shows significant gradient changes at the tumor infiltration edge, with a large total variation in the uncertainty gradient, while the lesion segmentation volume is relatively small, leading to a high ratio term. After substituting into the formula, the difficulty score falls into the advanced range. Another example is a typical lung nodule case, which contains only a single modality. The model has high classification confidence, low entropy value, and clear lesion boundaries, resulting in a small total variation in the uncertainty gradient. The lesion segmentation volume is moderate, and the ratio term is low, ultimately resulting in a beginner-level difficulty score. This formula transforms the intrinsic cognitive state of the artificial intelligence model into objective teaching indicators. Based on this, the system automatically organizes the case library to achieve a teaching sequence from easy to difficult.
[0049] In this invention, the teaching case grading module eliminates the need for manual difficulty labeling, reducing the cost of developing teaching resources. Simultaneously, the difficulty score directly correlates with the model's inference quality, ensuring that the teaching content aligns with the real challenges of AI-assisted diagnosis, thereby enhancing the clinical relevance of medical education and the effectiveness of cognitive training.
[0050] In one embodiment, such as Figure 2 As shown, a deep learning-based multimodal medical imaging teaching method is provided, including the following specific steps: S10: Receives medical image data from different imaging devices.
[0051] S20: The medical image data is mapped to the same anatomical coordinate system through differentiable affine transformation and non-rigid cascade transformation to generate aligned multimodal images.
[0052] S30: Process the aligned multimodal image using a task learning model to output a lesion segmentation mask, a lesion category probability vector, and a pixel-level uncertainty map.
[0053] S40: Simultaneously display the original multimodal image, the lesion segmentation mask, the heat map generated based on gradient class activation mapping, and the pixel-level uncertainty map, and receive user correction operations on the lesion segmentation results.
[0054] S50: Based on the user-corrected segmentation results, only update the parameters of the last two layers of the decoder in the task learning model.
[0055] S60: Calculate and classify the difficulty score of the teaching case based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
[0056] Specifically, the deep learning-based multimodal medical image interactive teaching method of the present invention includes the following steps. First, medical image data from different imaging devices are received. These different imaging devices include computed tomography (CT) scanners, magnetic resonance imaging (MRI) scanners, ultrasound imaging devices, and positron emission tomography (PET) scanners. The received medical image data is transmitted in DICOM 3.0 standard format and contains complete patient information and imaging parameters. Then, the medical image data is mapped to the same anatomical coordinate system through differentiable affine and non-rigid cascade transformations to generate aligned multimodal images. This process uses one modality as a reference modality, and the remaining modalities are sequentially corrected for overall pose through affine transformations, and then local deformations are compensated through non-rigid transformations based on displacement fields, ultimately achieving spatial consistency across all modalities at the voxel level, with a uniform voxel spacing of 1 mm × 1 mm × 1 mm. Next, a task learning model is used to process the aligned multimodal images, outputting a lesion segmentation mask, a lesion category probability vector, and a pixel-level uncertainty map. This model employs a shared encoder and multi-head decoder structure to simultaneously complete lesion localization, type discrimination, and confidence assessment. The system then synchronously displays the original multimodal image, the lesion segmentation mask, the heatmap generated based on gradient class activation mapping, and the pixel-level uncertainty map, and receives user corrections to the lesion segmentation results. The display uses a four-pane layout, supporting simultaneous slice browsing and timeline backtracking inference. Users can add or erase segmented regions using a drawing tool. The system records the corrected segmentation mask and triggers a subsequent update mechanism. Based on the user-corrected segmentation results, only the parameters of the last two layers of the decoder in the task learning model are updated. During the update process, all parameters of the encoder and decoder are frozen, and only single-step gradient descent is performed on the last two convolutional layers to generate a temporary model version specifically for the teaching session, without affecting the generalization ability of the main model. Finally, based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities, the difficulty score of the teaching case is calculated and graded. Entropy reflects classification uncertainty, total gradient variation is obtained by summing the absolute values of the first-order differences in three-dimensional space, lesion segmentation volume is the total number of voxels in the segmentation mask, modality number is the number of modal types in the input image, and the three are linearly combined according to preset weights to obtain a difficulty score, which is then divided into primary, intermediate or advanced teaching cases.
[0057] For example, the teacher imports a set of cases containing liver computed tomography (CT) scans and magnetic resonance imaging (MRI) images. The system automatically performs multimodal alignment, and the model outputs a liver tumor segmentation mask, a malignancy probability vector, and a high-uncertainty boundary map. Students observe in the four-pane interface that the model misclassifies some blood vessels as lesions. After removing these lesions using a correction tool, the system only fine-tunes the decoder's terminal parameters so that the region is no longer marked in subsequent slices. Simultaneously, because this case has two modalities, a moderate classification entropy value, and a high but small boundary uncertainty gradient, its difficulty score falls into the medium range, and it is included in the residency training database. This method achieves a complete teaching loop from data access, intelligent analysis, interactive correction to case grading. Students not only learn lesion identification knowledge but also understand the decision-making logic and limitations of artificial intelligence. Teachers can match teaching targets based on the automatic grading results, improving the relevance of training.
[0058] The deep learning-based multimodal interactive medical imaging teaching method described in this invention transforms artificial intelligence from a static demonstration tool into an interactive, evolving, and graded teaching carrier, significantly enhancing the depth, flexibility, and clinical authenticity of medical imaging education.
[0059] Optionally, the step of mapping medical image data to the same anatomical coordinate system includes: selecting one modal image as a reference modality; and sequentially performing affine transformation and displacement field-based non-rigid transformation on the remaining modal images to make their spatial coordinates consistent with the reference modality.
[0060] Optionally, the training loss function of the task learning model is composed of a weighted sum of segmentation loss, classification loss, and uncertainty loss; the uncertainty loss is the average of the product of pixel-level uncertainty value and the gradient magnitude of the segmentation mask at the corresponding position.
[0061] Optionally, the original multimodal fused image, semi-transparent lesion segmentation mask, gradient class activation map heatmap, and pixel-level uncertainty distribution map are presented through a four-pane interface; and a timeline control is provided to dynamically display the evolution sequence of feature maps during model inference.
[0062] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0063] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A deep learning-based multimodal medical imaging teaching system, characterized in that, include: The image access module is used to receive medical image data from different imaging devices; The spatial alignment module is used to map the medical image data to the same anatomical coordinate system through differentiable affine transformation and non-rigid cascade transformation, generating aligned multimodal images; The task learning model is used to process the aligned multimodal images and output lesion segmentation masks, lesion category probability vectors, and pixel-level uncertainty maps. The visualization module is used to simultaneously display the original multimodal image, the lesion segmentation mask, the heat map generated based on gradient class activation mapping, and the pixel-level uncertainty map, and supports users to correct the lesion segmentation results; The parameter update module is used to update only the parameters of the last two layers of the decoder in the task learning model based on the user-corrected segmentation results. The case classification module is used to calculate and classify the difficulty score of teaching cases based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
2. The deep learning-based multimodal medical imaging teaching system according to claim 1, characterized in that, The spatial alignment module uses one of the modal images as a reference modality, and the remaining modal images are aligned to the spatial coordinate system of the reference modality through a two-level transformation; the two-level transformation includes a first-level affine transformation and a second-level non-rigid transformation based on the displacement field.
3. The deep learning-based multimodal medical imaging teaching system according to claim 1, characterized in that, The task learning model adopts a shared encoder and multi-head decoder structure; the multi-head decoder includes a segmentation head, a classification head, and an uncertainty estimation head; the loss function of the uncertainty estimation head includes a product term of the pixel-level uncertainty map and the gradient magnitude of the segmentation mask.
4. The deep learning-based multimodal medical imaging teaching system according to claim 1, characterized in that, The visualization module uses a four-pane layout to simultaneously render the following content: the original multimodal fusion image; a semi-transparent overlay lesion segmentation mask; a gradient-class activation mapping heatmap; a pixel-level uncertainty distribution map; and supports backtracking the evolution of intermediate features during the model inference process via a timeline slider.
5. The deep learning-based multimodal medical imaging teaching system according to claim 1, characterized in that, After receiving the segmentation mask corrected by the user, the parameter update module freezes the encoder parameters of the task learning model and performs gradient descent updates on the parameters of the last two layers of the decoder only based on the segmentation loss function.
6. The deep learning-based multimodal medical imaging teaching system according to claim 1, characterized in that, The formula used by the case grading module to calculate the difficulty score of teaching cases is as follows: Difficulty score = α × entropy value + β × (total variation of uncertainty gradient / lesion segmentation volume) + γ × number of modes; Wherein, the entropy value is the information entropy of the lesion category probability vector; the total variation of the uncertainty gradient is the spatial gradient norm of the pixel-level uncertainty map; the lesion segmentation volume is the total number of voxels in the segmentation mask; the modality number is the number of modalities in the input image; and α, β, and γ are preset weight coefficients.
7. A multimodal medical image teaching method based on deep learning, characterized in that, include: Receives medical image data from different imaging devices; The medical image data is mapped to the same anatomical coordinate system through differentiable affine transformation and non-rigid cascade transformation to generate aligned multimodal images; The aligned multimodal image is processed using a task learning model to output a lesion segmentation mask, a lesion category probability vector, and a pixel-level uncertainty map. The system simultaneously displays the original multimodal image, the lesion segmentation mask, the heat map generated based on gradient class activation mapping, and the pixel-level uncertainty map, and receives user correction operations for the lesion segmentation results. Based on the user-corrected segmentation results, only the parameters of the last two layers of the decoder in the task learning model are updated; The difficulty score of the teaching case is calculated and graded based on the entropy value of the lesion category probability vector, the ratio of the total gradient variation of the pixel-level uncertainty map to the lesion segmentation volume, and the number of input modalities.
8. The deep learning-based multimodal medical image teaching method according to claim 7, characterized in that, The step of mapping medical image data to the same anatomical coordinate system includes: selecting one modal image as a reference modality; and sequentially performing affine transformation and displacement field-based non-rigid transformation on the remaining modal images to make their spatial coordinates consistent with the reference modality.
9. The deep learning-based multimodal medical image teaching method according to claim 7, characterized in that, The training loss function of the task learning model is composed of a weighted sum of segmentation loss, classification loss, and uncertainty loss; the uncertainty loss is the average of the product of pixel-level uncertainty value and the gradient magnitude of the segmentation mask at the corresponding position.
10. The deep learning-based multimodal medical image teaching method according to claim 7, characterized in that, The four-pane interface presents the original multimodal fusion image, the semi-transparent lesion segmentation mask, the gradient class activation mapping heatmap, and the pixel-level uncertainty distribution map, respectively; and provides a timeline control to dynamically display the evolution sequence of feature maps during the model inference process.