A lesion detection and labeling method and system for cardiac multi-modal medical images

By introducing the CIPEM module and the multi-scale lesion perception joint loss function, combined with a semi-supervised self-distillation learning framework, the problems of inconsistent annotation and insufficient cross-modal detection capabilities in intelligent cardiac image analysis are solved, achieving high-precision detection of small targets and low-contrast lesions.

CN122367971APending Publication Date: 2026-07-10GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610489761.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-14
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing technologies for intelligent analysis of cardiac images suffer from problems such as inconsistent annotation, time-consuming and labor-intensive processes, insufficient cross-modal detection capabilities, poor model generalization ability, and low lesion detection accuracy, especially for small targets and low-contrast lesions, which are difficult to detect effectively.

Method used

A cardiac image perception enhancement module (CIPEM) is introduced, which combines a multi-scale lesion perception joint loss function and a semi-supervised self-distillation learning framework. By dynamically and adaptively generating pseudo-labels, the model training process is optimized and the accuracy of lesion detection is improved.

Benefits of technology

It significantly improves the feature extraction capability and perception accuracy of complex lesions in cardiac imaging, especially for small targets and low-contrast lesions, and achieves high-precision cross-modal lesion detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367971A_ABST
    Figure CN122367971A_ABST
Patent Text Reader

Abstract

The application belongs to the field of medical image analysis, and provides a lesion detection and labeling method and system for cardiac multi-modal medical images, which comprises the following steps: acquiring multi-modal cardiac images and performing preprocessing to obtain preprocessed multi-modal cardiac images; based on the preprocessed multi-modal cardiac images, a pre-trained target detection model is used for lesion detection and labeling to obtain lesion positioning and labeling results; the application is trained in combination with a semi-supervised self-distillation labeling mechanism, and in the training process, a small amount of artificial labeling data and a large amount of high-quality pseudo labels generated through the self-distillation mechanism are used for joint optimization. At the same time, the designed dynamic prior box adaptive generation mechanism, cross-modal feature alignment module and progressive pseudo label refining strategy participate in the training and optimization of the model throughout the process, so that the generalization ability of the model on multi-modal data and the quality of the pseudo labels are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image analysis technology, specifically relating to a method and system for lesion detection and annotation in multimodal cardiac medical images. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] Heart disease is one of the leading causes of death worldwide, and early, accurate diagnosis is crucial for improving patient outcomes. Multimodal cardiac imaging, including CT, MRI, and ultrasound, is an important basis for clinical diagnosis, and precise detection and annotation of lesion areas in these images is the core foundation for computer-aided diagnosis. However, current technologies still face many technical bottlenecks in the field of intelligent cardiac image analysis.

[0004] Traditional cardiac lesion annotation relies heavily on the professional experience of radiologists. Faced with massive amounts of clinical image data, manual annotation is not only time-consuming and labor-intensive, but also often results in significant discrepancies between different physicians, making it difficult to guarantee consistency and repeatability. Some existing automated solutions employ a multi-stage preprocessing, segmentation, and post-processing pipeline, which is cumbersome, and errors between stages propagate and amplify at each stage, ultimately affecting detection accuracy. Due to significant differences in imaging principles, resolution, and noise characteristics among different modalities such as CT, MRI, and ultrasound, existing solutions typically require separate model development, training, and maintenance for each modality, resulting in high R&D costs and difficulty in achieving cross-modal knowledge transfer, leading to insufficient generalization ability of single models. Furthermore, the imaging characteristics of cardiac lesions themselves present inherent challenges for detection, such as small target size, low contrast with surrounding tissues, and imbalanced class distribution. General target detection models, such as the YOLO series, use a standard convolutional structure in their C2f module, resulting in a fixed receptive field and a lack of specialized enhancement mechanisms for small targets and low-contrast features, leading to a significant decrease in accuracy when processing cardiac images. Meanwhile, existing semi-supervised learning methods mostly use general contrastive learning loss or consistency regularization loss, which fail to be specifically designed for the multi-scale characteristics of lesions in cardiac images, the ambiguity of boundaries, and the uneven confidence of pseudo-labels, resulting in low quality of pseudo-labels and limited improvement in model performance. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a lesion detection and annotation method and system for multimodal cardiac medical images. By introducing a novel cardiac image perception enhancement module (CIPEM), this invention significantly improves the feature extraction capability and perception accuracy of the YOLOv11 model for complex lesions in cardiac images, especially for lesions with small targets, low contrast, and varied morphology.

[0006] According to some embodiments, the first aspect of the present invention provides a lesion detection and annotation method for cardiac multimodal medical imaging, employing the following technical solution: A lesion detection and annotation method for multimodal cardiac medical imaging includes: Acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images; Based on the preprocessed multimodal cardiac images, a pre-trained target detection model is used to detect and label lesions, and the lesion localization and labeling results are obtained. The training process of the target detection model includes: Obtain the original multimodal cardiac samples and preprocess them to obtain multimodal cardiac samples; Based on prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated. Based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate the initial pseudo-labels for the current round, and a multi-scale lesion perception joint loss function is used to optimize the initial pseudo-labels to obtain the pseudo-labels for the current round. The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-labels for the current round. Based on the joint training objective function, the target detection model is trained using the fused pseudo-labels of the current round until the maximum number of iterations is reached, resulting in a well-trained target detection model.

[0007] Furthermore, the target detection model adopts the YOLOv11 model, and the C2f module in the backbone network of the YOLOv11 model adopts the cardiac image perception enhancement module. The cardiac image perception enhancement module is composed of a multi-scale residual dense connection submodule, an adaptive channel recalibration submodule, and an edge enhancement submodule cascaded together. After processing the input feature map in sequence, the perception enhancement features are obtained.

[0008] Furthermore, the multi-scale residual dense connection submodule's processing includes: The input feature map is processed by four parallel dilated convolution branches. The results of the four dilated convolution branches are concatenated with the input feature map through channels and then convolved again to obtain the concatenated features. The concatenated features are residually connected to the input feature map to obtain residual dense features.

[0009] Furthermore, the adaptive channel recalibration submodule includes the following processing steps: After processing the residual concatenation features through a dual-path attention mechanism, we obtain global average pooling dimensionality reduction features and global max pooling dimensionality reduction features. The channel attention weights are obtained by adding the global average pooling dimensionality reduction features and the global max pooling dimensionality reduction features and then processing them with an activation function. Channel attention weights are used to recalibrate the residual dense features to obtain channel recalibrated features.

[0010] Furthermore, the edge enhancement submodule's processing includes: Edge features are extracted from the channel recalibration features based on the learnable Sobel horizontal kernel and the learnable Sobel vertical kernel, respectively, to obtain the horizontal edge response and the vertical edge response; Calculate edge amplitude based on horizontal and vertical edge responses; After convolution mapping of the edge amplitude, it is weighted and fused with the channel recalibration features to obtain the edge enhancement features.

[0011] Furthermore, the step of dynamically and adaptively generating dynamic prior boxes based on prior cardiac anatomy and multimodal cardiac samples includes: Based on the anatomical characteristics of cardiac images, four types of lesion regions are defined as prior cardiac anatomical structures. In the first iteration, clustering was performed based on cardiac multimodal samples to obtain an initial set of prior boxes; Calculate the matching degree between each cluster center in the initial prior box set and the prior cardiac anatomical structure; The matching degree between each cluster center and the prior heart anatomy is weighted and fused to generate the first round of dynamic prior boxes; In subsequent rounds, the predicted bounding boxes from the previous round and the dynamic prior bounding boxes from the previous round are fused to obtain the dynamic prior bounding boxes for the current round.

[0012] According to some embodiments, a second aspect of the present invention provides a lesion detection and annotation system for cardiac multimodal medical imaging, employing the following technical solution: A lesion detection and annotation system for multimodal cardiac medical imaging includes: The image processing module is configured to acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images. The lesion detection and annotation module is configured to detect and annotate lesions based on pre-processed multimodal cardiac images using a pre-trained target detection model, and obtain lesion localization and annotation results. The training process of the target detection model includes: Obtain the original multimodal cardiac samples and preprocess them to obtain multimodal cardiac samples; Based on prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated. Based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate the initial pseudo-labels for the current round, and a multi-scale lesion perception joint loss function is used to optimize the initial pseudo-labels to obtain the pseudo-labels for the current round. The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-labels for the current round. Based on the joint training objective function, the target detection model is trained using the fused pseudo-labels of the current round until the maximum number of iterations is reached, resulting in a well-trained target detection model.

[0013] According to some embodiments, a third aspect of the present invention provides a computer-readable storage medium.

[0014] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a lesion detection and annotation method for cardiac multimodal medical images as described in the first embodiment above.

[0015] According to some embodiments, a fourth aspect of the present invention provides a computer device.

[0016] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in the first embodiment above.

[0017] According to some embodiments, a fifth aspect of the present invention provides a computer program product or computer program.

[0018] A computer program product or computer program includes computer instructions stored in a computer-readable storage medium, wherein a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in a lesion detection and annotation method for cardiac multimodal medical images as described in the first embodiment above.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention significantly improves the YOLOv11 model's ability to extract features and improve the perception accuracy of complex lesions in cardiac images by introducing a novel cardiac image perception enhancement module (CIPEM), especially for lesions with small targets, low contrast, and varied morphology.

[0020] This invention aims to address the problem of scarce cardiac image annotation data. By designing an innovative semi-supervised self-distillation mechanism, high-quality pseudo-labels are generated using a large amount of unlabeled data. The pseudo-labels are then optimized by combining the multi-scale lesion perception joint loss function (MLJL) to improve the quality of the pseudo-labels and the learning efficiency of the model, thereby achieving high-precision detection with limited labeled data.

[0021] This invention draws upon the BYOL (Bootstrap Your Own Latent) self-supervised learning framework, enabling robust feature representations to be learned from unlabeled data without the need for negative samples through interactive learning between an online network and a target network. The inclusion of an EMA (Exponential Moving Average) teacher network ensures a more stable and high-quality pseudo-label generation process. Attached Figure Description

[0022] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0023] Figure 1 This is a schematic diagram of the training process of the overall method in an embodiment of the present invention; Figure 2 This is a schematic diagram of the architecture of the cardiac image perception enhancement module (CIPEM) in an embodiment of the present invention; Figure 3 This is a detailed architecture diagram of the multi-scale residual dense connection submodule (MR-Dense) in an embodiment of the present invention; Figure 4 This is a detailed architecture diagram of the Adaptive Channel Recalibration Submodule (ACR) in an embodiment of the present invention; Figure 5 This is a detailed architecture diagram of the edge enhancement submodule (EE) in an embodiment of the present invention; Figure 6 This is a schematic diagram of the composition of the multi-scale lesion sensing joint loss function (MLJL) in an embodiment of the present invention; Figure 7 This is an overall flowchart of the semi-supervised self-distillation learning framework in an embodiment of the present invention; Figure 8 This is a schematic diagram of the architecture of the cross-modal feature alignment module (CMFA) in an embodiment of the present invention; Figure 9 This is a flowchart of the dynamic prior box adaptive generation mechanism in an embodiment of the present invention. Detailed Implementation

[0024] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0025] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0026] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0027] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.

[0028] Example 1 This embodiment provides a method for lesion detection and annotation in multimodal cardiac medical imaging. This embodiment uses the application of this method to a server as an example for illustration. It is understood that this method can also be applied to a terminal, or to a system including a terminal, server, and system, and is implemented through interaction between the terminal and server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network servers, cloud communication, middleware services, domain name services, CDN security services, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein. In this embodiment, the method includes the following steps: Acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images; Based on the preprocessed multimodal cardiac images, a pre-trained target detection model is used to detect and label lesions, and the lesion localization and labeling results are obtained. The training process of the target detection model includes: (1) Obtain the original cardiac multimodal samples and preprocess them to obtain cardiac multimodal samples; (2) Based on the prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated; (3) Based on cardiac multimodal samples, the initial pseudo-labels for the current round are generated using a semi-supervised self-distillation learning framework, and the initial pseudo-labels are optimized using a multi-scale lesion perception joint loss function to obtain the pseudo-labels for the current round. (4) Weigh and fuse the pseudo-labels from the previous round with the pseudo-labels from the current round to obtain the pseudo-labels from the current round; (5) Based on the joint training objective function, train the target detection model using the fusion pseudo-label of the current round until the maximum number of iterations is reached, and obtain the trained target detection model.

[0029] This embodiment provides a method for lesion detection and annotation in multimodal cardiac medical imaging, the detailed process of which is as follows: Step S1: Acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images; Step S2: Based on the preprocessed multimodal cardiac images, lesion detection and annotation are performed using a pre-trained target detection model. The process is as follows: Step S2.1: Based on the preprocessed multimodal cardiac images, feature extraction is performed using the improved backbone network to output multi-scale feature maps for each modality. ; Specifically, this embodiment constructs a Cardiac ImagePerception Enhancement Module (CIPEM) to completely replace the C2f module (i.e., the C3k2 module, also called the C2f module) in the YOLOv11 model backbone network, resulting in an improved backbone network; such as Figure 2 As shown, the CIPEM module is specifically designed for the characteristics of cardiac images and consists of three cascaded sub-modules: the multi-scale residual dense connection sub-module (MR-Dense), the adaptive channel recalibration sub-module (ACR), and the edge enhancement sub-module (EE). Through a cascaded fusion mechanism, it achieves synergistic enhancement to obtain the improved YOLOv11 model, which serves as the target detection model in this embodiment.

[0030] like Figure 3 As shown, the Multi-scale Residual Dense Connection Module (MR-Dense) aims to address the challenges of large size and varied morphology of cardiac lesions. This module employs dilated convolutions with different dilation rates in parallel to capture multi-scale features and achieves feature reuse through dense connections. Its architecture design is as follows: Let the input feature map be The MR-Dense submodule contains four parallel dilated convolution branches with dilation rates of [missing information]. The corresponding receptive fields are respectively , , , The calculation formula is as follows:

[0031] in, Indicates the expansion rate of Dilated convolution, BN represents batch normalization, and ReLU is the activation function. Indicates the first The output of each dilated convolution branch.

[0032] A dense connection mechanism is used to concatenate the four branches with the input feature map through channels before convolution, resulting in concatenated features. The formula is as follows:

[0033] Here, Concat represents the concatenation operation at the channel level. Used for channel number compression and feature fusion.

[0034] The concatenated features are residually concatenated with the input feature map to obtain residual dense features. The formula for residual connection is as follows:

[0035] The residual connection consists of four branches, as shown in Table 1.

[0036] Table 1 Residual Connection Structure

[0037] Based on residual characteristics As the output of the multi-scale residual dense connection submodule.

[0038] like Figure 4 As shown, the Adaptive Channel Recalibration Module (ACR) dynamically learns the dependencies between channels, adaptively enhances feature channels related to lesions, and suppresses background noise channels. Architecture design: The ACR submodule employs a dual-path attention mechanism, simultaneously considering statistical information extracted by global average pooling and global max pooling, and learns channel weights through a shared multilayer perceptron.

[0039] The formulas for global average pooling and global max pooling are as follows:

[0040]

[0041] in, It is a global average pooling feature. It is a global max pooling feature. and These represent the global average pooling function and the global max pooling function, respectively.

[0042] By utilizing a shared multilayer perceptron, dimensionality reduction of a shared MLP is achieved, yielding pooled dimensionality reduction results for two paths, as shown in the following formula:

[0043]

[0044] in, It is a global average pooling dimensionality reduction feature. It is a global max pooling dimensionality reduction feature. and To share the weight matrix, This is the compression ratio (the default setting is 16).

[0045] After summing the pooling and dimensionality reduction results of the two paths, the channel attention weights are obtained by applying an activation function, as shown in the following formula:

[0046] in, This represents the Sigmoid activation function.

[0047] Utilizing channel attention weights for dense residual features of the input Perform feature recalibration to obtain channel recalibration features. ,as follows:

[0048] in, This indicates multiplication by channel.

[0049] like Figure 5 As shown, the Edge Enhancement Module (EE) is specifically designed to address the problem of blurred boundaries and low contrast with surrounding tissues in cardiac lesions. This submodule uses a learnable Sobel operator to enhance the lesion boundary features, and its architecture is as follows: Building a learnable Sobel horizontal kernel and learnable Sobel vertical kernel ,as follows:

[0050]

[0051] in, and For learnable Sobel level kernels and learnable Sobel vertical kernel The Middle Line 1 The learnable perturbation parameters of the column elements are initialized to 0 and adaptively learned to obtain the optimal edge detection kernel through backpropagation.

[0052] Based on learnable Sobel level kernel and learnable Sobel vertical kernel Recalibrate the input channels separately Edge feature extraction is performed to obtain the horizontal edge response. and vertical edge response The formula is as follows:

[0053]

[0054] Based on horizontal edge response and vertical edge response Calculate edge amplitude The formula is as follows:

[0055] After convolution mapping of edge magnitudes, the features are recalibrated with the input channels. Weighted fusion is performed to obtain edge enhancement features. The formula is as follows:

[0056] in, For learnable fusion weights, Used to map edge magnitudes (the edge magnitude map of the image, representing the edge intensity at each location) to... Same channel dimension.

[0057] The cascaded fusion mechanism of the CIPEM module yields perceptual enhancement features. The complete processing flow is as follows:

[0058] Perceptual augmentation features and input features are used to perform residual augmentation to obtain perceptual residual augmentation features. ,as follows:

[0059] in, This is a learnable residual scaling factor, initialized to 0.1, used to stabilize gradient flow in the early stages of training.

[0060] The architecture of the cardiac image perception enhancement module (CIPEM) is shown in Table 2.

[0061] Table 2. Architecture Composition of Cardiac Image Perception Enhancement Module

[0062] This embodiment introduces a novel Cardiac Image Perception Enhancement Module (CIPEM), significantly improving the YOLOv11 model's feature extraction capability and perception accuracy for complex lesions in cardiac images, especially for small targets, low-contrast lesions, and morphologically variable lesions. The CIPEM module consists of three cascaded core components: a multi-scale residual dense connection submodule (MR-Dense), an adaptive channel recalibration submodule (ACR), and an edge enhancement submodule (EE). It completely replaces the original C2f module in YOLOv11, forming the target detection model in this embodiment, thereby achieving more refined feature representation without increasing the computational burden excessively.

[0063] Step S2.2: Based on the multi-scale feature maps of each modality, use the neck network to perform feature fusion and enhancement to obtain the fused features corresponding to each scale; Step S2.3: The detection head processes the fusion features corresponding to each scale and sends the fusion features corresponding to each scale into the YOLOv11 detection head for lesion localization (boundary box regression) and classification prediction.

[0064] Step S2.4: Post-processing and output Non-maximum suppression (NMS) removes redundant boxes and outputs the final lesion detection results (location, category, confidence).

[0065] The final output of lesion detection results includes the location coordinates, category label, and confidence score of each lesion, which can be directly used for the localization, classification, and automatic annotation of cardiac lesions.

[0066] Among them, such as Figure 1 As shown, the training process of the target detection model includes: (1) Obtain the original cardiac multimodal samples and preprocess them to obtain cardiac multimodal samples; Acquire cardiac images in three modalities: CT, MRI, and ultrasound, as raw cardiac samples; and preprocess the raw cardiac samples to obtain multimodal cardiac samples; the preprocessing here includes standardized preprocessing operations such as size adjustment and normalization.

[0067] This embodiment combines a semi-supervised self-distillation annotation mechanism based on the MLJL loss function for training. During training, a small amount of manually labeled data and a large number of high-quality pseudo-labels generated through the self-distillation mechanism are used for joint optimization. Simultaneously, a designed dynamic prior box adaptive generation mechanism, a CMFA cross-modal feature alignment module, and a progressive pseudo-label refinement strategy are all involved in the model's training and optimization throughout the process, ensuring the model's generalization ability on multimodal data and the quality of the pseudo-labels.

[0068] (2) Based on the prior knowledge of cardiac anatomy and cardiac multimodal samples, dynamic prior boxes are dynamically and adaptively generated. In the first iteration, the initial prior boxes are generated by weighted fusion of K-means++ clustering results and prior knowledge of cardiac anatomy. In subsequent iterations, the predicted boxes are dynamically updated by combining the mean of the predicted boxes from the previous iteration. The process is as follows: This embodiment designs a dynamic prior box adaptive generation mechanism based on prior knowledge of cardiac anatomy, such as... Figure 9 As shown, instead of traditional fixed prior boxes or simple K-means clustering methods, the traditional YOLO series models use preset fixed prior boxes, which are difficult to adapt to the characteristics of the variable and large size difference of cardiac lesions. This mechanism improves the recall rate and localization accuracy by integrating prior knowledge of cardiac anatomy and dynamically generating prior boxes that better match the current image content.

[0069] Based on the anatomical characteristics of cardiac images, four types of lesion regions are defined as prior cardiac anatomical structures. Specifically, based on the anatomical characteristics of cardiac images, four types of lesion area size distribution priors are defined, as shown in Table 3.

[0070] Table 3 Prior knowledge of cardiac anatomy

[0071] In the first iteration, clustering was performed based on cardiac multimodal samples to obtain an initial set of prior boxes; Specifically, the width and height information of all lesion bounding boxes are extracted from the labeled data of cardiac multimodal samples to form a two-dimensional dataset. ,in The total number of bounding boxes is given; the K-means++ algorithm is used to cluster the dataset to obtain... 1 cluster center as the initial set of prior boxes:

[0072] in, Indicates the first The width and height of the initial prior bounding box. It can be set to 9, but is not limited to, to cover cardiac lesions of different scales.

[0073] The clustering objective function of the K-means++ algorithm is:

[0074] Among them, distance metric We use IoU-based distance to make the clustering results more consistent with the bounding box matching criteria in object detection.

[0075] Calculate the matching degree between each cluster center in the initial prior box set and the prior cardiac anatomical structure. The formula is as follows:

[0076] in, For the first Prior dimensions of cardiac anatomy for lesion-like structures. The Gaussian kernel bandwidth parameter controls the decay rate of the fit between the prior bounding box and the anatomical structure. It can be set to an empirical value or determined through cross-validation. It is the first in the initial prior box set Cluster centers, , for and The square of the Euclidean distance between them.

[0077] The matching degree between each cluster center and the prior knowledge of the cardiac anatomy is weighted and fused to generate the first round of dynamic prior boxes. The formula is as follows:

[0078] in, The fusion coefficient is... Weights for lesion types.

[0079] In subsequent rounds, the average predicted bounding boxes from the previous round and the dynamic prior bounding boxes from the previous round are fused to obtain the dynamic prior bounding boxes for the current round, specifically: During the training process, every Each epoch dynamically updates the prior bounding box based on the prediction results:

[0080] in, For the first Mean value of the predicted bounding box for the target lesion. For momentum parameters, , It was the previous round. No. Dynamic prior bounding boxes for target lesions. It is the current round number Dynamic prior boxes for target lesions. This update mechanism ensures that the prior boxes evolve smoothly during training, avoiding training instability caused by drastic fluctuations.

[0081] By fusing K-means++ adaptive clustering with prior knowledge of cardiac anatomy, the size of the prior bounding box is made to better match the actual distribution of cardiac lesions, thus improving the recall rate of small target detection. The EMA dynamic update mechanism enables the prior bounding box to evolve smoothly with training, adapting to changes in lesion distribution.

[0082] (3) Based on cardiac multimodal samples, the initial pseudo-labels for the current round are generated using a semi-supervised self-distillation learning framework, and the initial pseudo-labels are optimized using a multi-scale lesion perception joint loss function to obtain the pseudo-labels for the current round. To address the problem of scarce cardiac imaging labeled data, an innovative semi-supervised self-distillation mechanism is designed to generate high-quality pseudo-labels using a large amount of unlabeled data. This mechanism is then combined with a multi-scale lesion perception joint loss function (MLJL) to optimize the quality of the pseudo-labels and the learning efficiency of the model, thereby achieving high-precision detection with limited labeled data.

[0083] First, based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate initial pseudo-labels for the current round; in the first iteration, the pseudo-labels of the current round are directly used as fused pseudo-labels. Specifically, the complete process of the semi-supervised self-distillation learning framework, such as... Figure 7 As shown, a BYOL framework consisting of an online network (student network) and a target network (teacher network) is constructed. Two different random data augmentations are applied to the same batch of unlabeled cardiac images, generating two distinct views. The online network processes one view and outputs a prediction, while the target network processes the other view and outputs a target representation (pseudo-label). It can be understood that the online network and target network described in this embodiment are the target detection model to be trained, i.e., the improved YOLOv11 model. This step aims to address the scarcity of labeled cardiac imaging data. An innovative semi-supervised self-distillation mechanism is designed to generate high-quality pseudo-labels from a large amount of unlabeled data. This is combined with a multi-scale lesion perception joint loss function (MLJL) to optimize the quality of the pseudo-labels and the model's learning efficiency, thereby achieving high-precision detection with limited labeled data. The BYOL (Bootstrap Your Own Latent) teacher network is constructed, drawing inspiration from the BYOL self-supervised learning framework. Through interactive learning between the online network and the target network, robust feature representations can be learned from unlabeled data without negative samples. The integration of an EMA (Exponential Moving Average) teacher network ensures a more stable and high-quality pseudo-label generation process.

[0084] Secondly, the initial pseudo-labels are optimized using a multi-scale lesion sensing joint loss function to obtain the pseudo-labels for the current round; Specifically, addressing the issues of large lesion size variations, blurred boundaries, and uneven pseudo-label quality in cardiac imaging, a semi-supervised self-distillation annotation mechanism based on the Multi-scale Lesion-aware Joint Loss (MLJL) is introduced. This mechanism, combined with the BYOL framework and the EMA teacher network, achieves high-quality pseudo-label generation. The Multi-scale Lesion-aware Joint Loss (MLJL) optimizes the semi-supervised self-distillation learning process, comprehensively improving the model's detection performance for multi-scale lesions and the quality of pseudo-label generation. Figure 6 As shown, the MLJL loss function consists of three innovation loss terms: scale-adaptive contrast loss (SAC Loss), boundary-aware consistency loss (BAC Loss), and confidence calibration loss (CC Loss).

[0085] The joint loss function for multi-scale lesion sensing is as follows:

[0086] in, , , The weighting coefficients for each loss term are set to [default value]. , , .

[0087] Scale-Adaptive Contrastive Loss (SAC Loss) addresses the large differences in the size of cardiac lesions by dynamically adjusting the weights of contrastive learning according to the target scale, ensuring that the model receives sufficient supervision signals at all scales.

[0088] Scale weight calculation, assuming the area of ​​the target bounding box is... , define the first Batch Scale Weighting Function ,as follows:

[0089] This weighting function gives higher weight to smaller targets, balancing the learning difficulty of targets at different scales.

[0090] By comparing the learning objectives, a scale-adaptive contrastive loss function is obtained. The formula is as follows:

[0091] in, Features extracted from student networks Positive sample features extracted from the teacher network. Represents cosine similarity. This is the temperature parameter (default setting is 0.07). For batch size, The first Batch and No. batch.

[0092] Existing contrastive learning loss treats all samples equally, while SAC Loss, through a scale weighting mechanism, explicitly assigns higher learning priority to small targets, effectively solving the problem that small lesions in cardiac images are easily overlooked.

[0093] Boundary-Aware Consistency Loss (BAC Loss) enhances the accuracy of pseudo-label boundaries, requiring a high degree of consistency between the boundaries predicted by the student network and the pseudo-label boundaries generated by the teacher network. Boundary extraction involves extracting the boundaries of the predicted bounding boxes, defining the boundary regions as the bounding boxes. Expanding both internally and externally The formula for the annular region of a pixel is as follows:

[0094] in, Point Distance to the boundary of the bounding box, It is the first Each pixel.

[0095] Then, the boundary-aware consistency loss function ,as follows:

[0096] in, and These are feature maps of the student network and the teacher network in the boundary region, respectively. and For the predicted bounding box and the pseudo-labeled bounding box, GIoU is the generalized intersection-union ratio.

[0097] Traditional consistency loss only constrains the consistency of overall predictions, while BAC Loss specifically strengthens the constraints on the boundary region, effectively solving the problem of inaccurate localization caused by the fuzzy boundary of cardiac lesions.

[0098] Confidence Calibration Loss (CC Loss) dynamically calibrates the confidence distribution of pseudo-labels, preventing the teacher network from generating overconfident or underconfident pseudo-labels. Expected Calibration Error (ECE) optimization divides the predicted confidence intervals into... If there are bins of equal width, then the confidence calibration loss function... definition:

[0099] in, To fall into the first A sample set of bins, and The first The average accuracy and average confidence of samples within each bin.

[0100] Temperature scaling calibration introduces learnable temperature parameters. The teacher's network output is calibrated as follows:

[0101] Dynamic threshold adjustment, based on the calibrated confidence distribution, dynamically adjusts the pseudo-label screening threshold as follows:

[0102] in, This is the initial threshold (default 0.5). To adjust the coefficient, This is the calibration loss from the previous round.

[0103] Existing methods mostly use fixed thresholds to screen pseudo-labels, while CC Loss uses a dynamic calibration mechanism to adaptively adjust the threshold according to the model training state, ensuring that high-quality pseudo-labels can be selected at each stage of training. The composition of the multi-scale lesion perception joint loss function is shown in Table 4.

[0104] Table 4. Composition of MLJL Loss Function

[0105] (4) The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-label cardiac multimodal samples from the current round; The progressive pseudo-label refinement strategy improves pseudo-label quality through multiple iterations. Employing a curriculum learning approach, it gradually expands the training set from simple to difficult samples. In semi-supervised learning, the quality of pseudo-labels directly impacts model performance. This strategy, through curriculum learning and a multi-round fusion mechanism, progressively improves the accuracy and reliability of pseudo-labels, thereby more effectively utilizing unlabeled data. The refinement process is shown in Table 5.

[0106] Table 5 Refining Process

[0107] Weighted fusion of pseudo-labels generated from the same sample in different rounds:

[0108] in, It integrates pseudo-labels and weights. Dynamic allocation based on the validation set performance of that round:

[0109] in, It is the summation index in the summation formula, traversing from 1 to... All rounds are used to calculate the denominator (the sum of mAP for all rounds). .

[0110] (5) Based on the joint training objective function, train the target detection model using the fused pseudo-labels of the current round until the maximum number of iterations is reached, and obtain the trained target detection model; EMA teacher network updates: The parameters of the teacher network are updated from the student network using an exponential moving average.

[0111] in, For momentum parameters, It is the first Parameters of the EMA teacher network during the step. It is the first Parameters of the online network (student network) during walks.

[0112] The joint training objective function is as follows:

[0113] in, The standard detection loss includes classification loss and regression loss. It is a multi-scale lesion sensing joint loss function. These are the weighting coefficients (hyperparameters) of the CMFA loss, used to balance the contribution of cross-modal feature alignment loss to the total loss. It is a cross-modal feature alignment loss.

[0114] For the standard detection loss: Based on the intersection-union ratio (IU) of the dynamic prior boxes and the ground truth boxes, positive and negative samples are assigned according to the IU, and the standard detection loss is constructed by using the offset between the dynamic prior boxes of the positive samples and the corresponding ground truth boxes. Among them, standard detection loss It consists of classification loss and regression loss:

[0115] in, The regression loss is used (this invention uses CIoU Loss to improve positioning accuracy). For classification loss (Focal Loss is used to alleviate the imbalance between positive and negative samples). It is the number of positive samples. It is the number of negative samples. and The first The predicted offset and the encoded true offset of a positive sample. and The first The predicted class probability and true label of each sample are calculated. A relative offset encoding method is used to unify the learning difficulty of targets at different scales, enhancing the model's scale invariance. A combination of CIoU Loss and Focal Loss is employed to effectively improve localization accuracy and alleviate the problem of imbalanced positive and negative samples.

[0116] For cross-modal feature alignment loss: feature extraction is performed on cardiac multimodal samples respectively, and adversarial loss, maximum mean difference loss and semantic consistency loss are determined based on the features of different modalities. The adversarial loss, maximum mean difference loss and semantic consistency loss are weighted and fused to obtain cross-modal feature alignment loss. like Figure 8 As shown, the Cross-Modal Feature Alignment Module (CMFA) achieves explicit alignment of the feature spaces of CT, MRI, and ultrasound, eliminates modal gaps, realizes effective fusion of multimodal information, and enhances the model's cross-modal generalization ability.

[0117] Modality-specific encoders set up independent shallow encoders for each mode to extract modal features:

[0118] Based on cross-modal alignment loss, including: Domain adversarial training, introducing a modality discriminator Adversarial training makes the features of different modalities indistinguishable in the shared space, resulting in adversarial loss, calculated as follows:

[0119] Maximum mean difference (MMD) alignment, the formula is as follows:

[0120] in, For kernel mapping functions, Regenerating Hilbert space, This represents the number of samples in the source domain. This represents the number of samples in the target domain. It represents the first in the source domain. The feature vector of each sample It represents the first in the target domain. The feature vector of each sample.

[0121] Semantic consistency constraints, the formula is as follows:

[0122] in, For modality Next category The prototype feature vector.

[0123] The cross-modal feature alignment loss (CMFA total loss) is calculated as follows:

[0124] in, This is the weighting coefficient (balancing factor) of the MMD loss, used to control the proportion of the MMD alignment loss in the total loss. It is the weighting coefficient (balancing factor) of semantic consistency loss, used to control the proportion of semantic consistency constraints in the total loss.

[0125] Multimodal unified modeling and lightweight inference, learnable modality tokens, assigning a learnable modality-specific vector to each imaging modality (CT, MRI, ultrasound):

[0126] At the model input end, the token is concatenated with the image patch embedding sequence and then input into the Transformer layer to achieve cross-modal adaptive inference.

[0127] Structured sparse pruning employs an L1-norm-based pruning criterion to evaluate the importance scores of all convolutional kernels and attention heads in the model, removing redundant parameters with minimal contributions. The pruning ratio is set to 30%-50%, achieving model compression while maintaining an accuracy loss of less than 1%.

[0128] Post-training quantization (PTQ) converts all weights and activations in the model from FP32 to INT8 format, with the option to further reduce it to INT4. After pruning and quantization compression, the model achieves a single-frame inference latency of approximately 9 milliseconds on edge computing devices.

[0129] Based on the aforementioned enhanced model and innovative mechanism, high-precision localization, classification, and automatic annotation of cardiac lesions are achieved. However, single-round pseudo-labels inherently contain noise and uncertainty, especially for small lesions in cardiac imaging (such as microcalcifications and early plaques), where a single prediction may result in missed or false detections. Through multi-round iterative fusion, the impact of these noises can be effectively reduced, making the final pseudo-labels more reliable.

[0130] The trained target detection model can receive cardiac images from any modality, such as CT, MRI, or ultrasound, as input. The CIPEM module, a cardiac image perception enhancement module, extracts depth features and enhances perception from the input images. Combined with the YOLOv11 detection head, the model can accurately output the bounding box coordinates (localization) and corresponding lesion categories (classification) of cardiac lesions, such as myocardial infarction, valvular disease, and tumors. The CMFA module, a cross-modal feature alignment module, ensures that the model can effectively fuse modal information when processing images from different modalities, maintaining consistently high accuracy performance.

[0131] Automatic annotation: After locating and classifying lesions, the model can directly generate annotation files conforming to standard formats such as DICOM or COCO. These annotation files contain information such as lesion category, location, and confidence level, which can be used to assist doctors in diagnosis or as a data foundation for subsequent medical research. The introduction of a semi-supervised self-distillation mechanism enables the model to automatically annotate large-scale image data at extremely low annotation costs, greatly improving the efficiency of clinical work.

[0132] The resulting intelligent detection and annotation method can be deployed on medical imaging workstations, cloud platforms, or edge computing devices. Through lightweight inference techniques (such as model pruning and quantization), it ensures that the model can achieve real-time or near-real-time inference speeds even in resource-constrained environments, meeting the immediate needs of clinical diagnosis. This system can provide strong technical support for the early screening, diagnosis, and treatment assessment of heart diseases.

[0133] Example 2 This embodiment provides a lesion detection and annotation system for multimodal cardiac medical imaging, including: The image processing module is configured to acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images. The lesion detection and annotation module is configured to detect and annotate lesions based on pre-processed multimodal cardiac images using a pre-trained target detection model, and obtain lesion localization and annotation results. The training process of the target detection model includes: Obtain the original multimodal cardiac samples and preprocess them to obtain multimodal cardiac samples; Based on prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated. Based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate the initial pseudo-labels for the current round, and a multi-scale lesion perception joint loss function is used to optimize the initial pseudo-labels to obtain the pseudo-labels for the current round. The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-labels for the current round. Based on the joint training objective function, the target detection model is trained using the fused pseudo-labels of the current round until the maximum number of iterations is reached, resulting in a well-trained target detection model.

[0134] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1 above. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0135] The descriptions of each embodiment in the above embodiments have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0136] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and the division of modules described above is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed.

[0137] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in Embodiment 1 above.

[0138] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in Embodiment 1 above.

[0139] Example 5 This embodiment provides a computer program product or computer program, including computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in the lesion detection and annotation method for cardiac multimodal medical images described in Embodiment 1 above.

[0140] Those skilled in the art will understand that embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0141] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0142] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0143] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0144] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0145] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.

Claims

1. A method for lesion detection and annotation in multimodal cardiac medical imaging, characterized in that, include: Acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images; Based on the preprocessed multimodal cardiac images, a pre-trained target detection model is used to detect and label lesions, and the lesion localization and labeling results are obtained. The training process of the target detection model includes: Obtain the original multimodal cardiac samples and preprocess them to obtain multimodal cardiac samples; Based on prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated. Based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate the initial pseudo-labels for the current round, and a multi-scale lesion perception joint loss function is used to optimize the initial pseudo-labels to obtain the pseudo-labels for the current round. The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-labels for the current round. Based on the joint training objective function, the target detection model is trained using the fused pseudo-labels of the current round until the maximum number of iterations is reached, resulting in a well-trained target detection model.

2. The lesion detection and annotation method for cardiac multimodal medical imaging as described in claim 1, characterized in that, The target detection model adopts the YOLOv11 model, and the C2f module in the backbone network of the YOLOv11 model adopts the cardiac image perception enhancement module. The cardiac image perception enhancement module is composed of a multi-scale residual dense connection submodule, an adaptive channel recalibration submodule, and an edge enhancement submodule cascaded together. After processing the input feature map in sequence, the perception enhancement features are obtained.

3. The lesion detection and annotation method for cardiac multimodal medical imaging as described in claim 2, characterized in that, The multi-scale residual dense connection submodule's processing includes: The input feature map is processed by four parallel dilated convolution branches. The results of the four dilated convolution branches are concatenated with the input feature map through channels and then convolved again to obtain the concatenated features. The concatenated features are residually connected to the input feature map to obtain residual dense features.

4. The lesion detection and annotation method for cardiac multimodal medical imaging as described in claim 2, characterized in that, The adaptive channel recalibration submodule, the processing procedure includes: After processing the residual concatenation features through a dual-path attention mechanism, we obtain global average pooling dimensionality reduction features and global max pooling dimensionality reduction features. The channel attention weights are obtained by adding the global average pooling dimensionality reduction features and the global max pooling dimensionality reduction features and then processing them with an activation function. Channel attention weights are used to recalibrate the residual dense features to obtain channel recalibrated features.

5. The lesion detection and annotation method for cardiac multimodal medical imaging as described in claim 2, characterized in that, The edge enhancement submodule's processing includes: Edge features are extracted from the channel recalibration features based on the learnable Sobel horizontal kernel and the learnable Sobel vertical kernel, respectively, to obtain the horizontal edge response and the vertical edge response; Calculate edge amplitude based on horizontal and vertical edge responses; After convolution mapping of the edge amplitude, it is weighted and fused with the channel recalibration features to obtain the edge enhancement features.

6. The lesion detection and annotation method for cardiac multimodal medical imaging as described in claim 1, characterized in that, The process of dynamically and adaptively generating dynamic prior boxes based on prior knowledge of cardiac anatomy and multimodal cardiac samples includes: Based on the anatomical characteristics of cardiac images, four types of lesion regions are defined as prior cardiac anatomical structures. In the first iteration, clustering was performed based on cardiac multimodal samples to obtain an initial set of prior boxes; Calculate the matching degree between each cluster center in the initial prior box set and the prior cardiac anatomical structure; The matching degree between each cluster center and the prior heart anatomy is weighted and fused to generate the first round of dynamic prior boxes; In subsequent rounds, the predicted bounding boxes from the previous round and the dynamic prior bounding boxes from the previous round are fused to obtain the dynamic prior bounding boxes for the current round.

7. A lesion detection and annotation system for multimodal cardiac medical imaging, characterized in that, include: The image processing module is configured to acquire multimodal cardiac images and perform preprocessing to obtain preprocessed multimodal cardiac images. The lesion detection and annotation module is configured to detect and annotate lesions based on pre-processed multimodal cardiac images using a pre-trained target detection model, and obtain lesion localization and annotation results. The training process of the target detection model includes: Obtain the original multimodal cardiac samples and preprocess them to obtain multimodal cardiac samples; Based on prior knowledge of cardiac anatomy and multimodal cardiac samples, dynamic prior boxes are dynamically and adaptively generated. Based on cardiac multimodal samples, a semi-supervised self-distillation learning framework is used to generate the initial pseudo-labels for the current round, and a multi-scale lesion perception joint loss function is used to optimize the initial pseudo-labels to obtain the pseudo-labels for the current round. The fused pseudo-labels from the previous round are weighted and fused with the pseudo-labels from the current round to obtain the fused pseudo-labels for the current round. Based on the joint training objective function, the target detection model is trained using the fused pseudo-labels of the current round until the maximum number of iterations is reached, resulting in a well-trained target detection model.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in any one of claims 1-6.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps in the lesion detection and annotation method for cardiac multimodal medical images as described in any one of claims 1-6.