Method and system for segmenting and classifying space occupying lesion area in cardiac cavity based on ultrasonic image

By using the dual-branch architecture of the CardioMassNet model, automated segmentation and classification of intracardiac space-occupying lesions were achieved, solving the problem of traditional ultrasound diagnosis relying on experience, improving segmentation accuracy and classification consistency, and reducing computational costs and risks.

CN121707919APending Publication Date: 2026-03-20TONGJI HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511653520.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

Traditional echocardiography diagnosis of intracardiac space-occupying lesions relies on operator experience, is easily affected by image quality, and lacks automated tools, resulting in inaccurate lesion segmentation and poor classification consistency.

Method used

We employ the CardioMassNet model based on the U-net architecture, combining an encoder-decoder architecture and a global max pooling layer to construct a dual-branch output module for lesion region segmentation and classification. We use the training set and validation set for model training and evaluation to achieve automated segmentation and classification.

Benefits of technology

It effectively narrows the interpretation differences between different experience levels, reduces computational costs and overfitting risks, enhances the ability to discriminate complex patterns, and provides intelligent auxiliary analysis tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121707919A_ABST
    Figure CN121707919A_ABST
Patent Text Reader

Abstract

The invention provides a segmentation and classification method and system for a space-occupying lesion area in a cardiac cavity based on an ultrasonic image, and the method comprises the steps: embedding a global maximum pooling layer at the output end of each stage of a U-Net encoder-decoder, splicing multi-scale feature vectors, inputting the spliced multi-scale feature vectors into a classification branch, keeping the tail end feature of the decoder to reach a segmentation branch, and carrying out the segmentation of the space-occupying lesion area in the U-Net encoder-decoder. A dual-task architecture with a shared trunk and independent output is formed; the calculation cost and the over-fitting risk are reduced through parameter multiplexing and task regularization, and the capability of discriminating complex modes is enhanced through cross-level feature fusion. Through multi-center heterogeneous data and time sequence verification training, the method has strong robustness for differences of ultrasonic image acquisition equipment and image quality fluctuation, can be seamlessly integrated to an existing ultrasonic workstation, assists image analysis through visual masks and quantitative probabilities, effectively narrows interpretation differences among different experience levels, and improves the interpretation accuracy. The risk of traditional enhanced examination is avoided, and an intelligent auxiliary technical scheme capable of being applied in a large scale is provided for clinic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical image analysis technology, and in particular to a method and system for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images. Background Technology

[0002] Intracardiac masses (ICMs) are a highly heterogeneous group of cardiac diseases, including thrombi, benign and malignant tumors, and vegetations. They can be discovered incidentally, from asymptomatic cases to fatal embolisms, heart failure, or malignant arrhythmias. These lesions are often discovered incidentally during echocardiography or CT scans of the chest and abdomen for non-cardiac indications. Although pathological biopsy is the gold standard for diagnosis, ICM biopsy requires a high level of operator experience, and even with advancements in catheter and guidewire techniques improving safety, sampling difficulties and potential complications remain for rare or unusually located lesions. In most clinical situations, accurately differentiating thrombi from tumors on imaging is crucial for treatment decisions.

[0003] Transthoracic echocardiography (TTE) is considered the preferred initial screening tool for intracranial lesions (ICMs) due to its non-invasive, real-time, and bedside operation characteristics. It can assess lesion size, location, activity, and hemodynamic effects, and is more suitable for dynamic monitoring than CT and cardiac magnetic resonance imaging (CMR). However, the accuracy and efficiency of traditional ultrasound diagnosis still rely on operator experience. Contrast-enhanced ultrasound (CEUS) can differentiate between thrombotic and non-thrombotic lesions based on perfusion in the lesion area. However, benign tumors (such as myxomas) typically have poor blood supply and exhibit hypoperfusion. Furthermore, the risk of contrast agent allergy and post-processing requirements limit its widespread adoption.

[0004] Therefore, there is a need to propose a system and method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, which can automatically segment and classify thrombotic and non-thrombotic lesions on two-dimensional echocardiography, providing a new analytical tool to reduce the potential fatal risks to patients. Summary of the Invention

[0005] In view of this, the present invention provides a system and method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, in order to solve the technical problems of traditional echocardiography diagnosis of intracardiac space-occupying lesions, which rely too much on operator experience, are easily affected by image quality, and lack automated tools, resulting in inaccurate segmentation and poor classification consistency of ultrasound-based ICM lesion regions.

[0006] To achieve the above-mentioned technical objectives, the present invention adopts the following technical solution:

[0007] On the one hand, the present invention provides a method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, including:

[0008] The raw echocardiogram images were acquired, labeled, and preprocessed to obtain an image dataset, which was then divided into a training set and a validation set.

[0009] The CardioMassNet model is constructed based on the U-net architecture. The model employs an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max-pooling layers are configured at the output of each feature extraction stage in the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The model's output is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch converts the feature map output by the decoder into a pixel-level prediction mask to segment intracardiac lesion regions. The classification branch inputs the multi-scale global feature vector formed by concatenating the global pooling feature vectors from each stage into a fully connected layer to predict at least two class labels for the lesion region.

[0010] The CardioMassNet model is trained using the training set, and the parameters are saved after training to obtain the trained CardioMassNet model.

[0011] The performance of the trained CardioMassNet model was evaluated using a validation set. After the evaluation, the CardioMassNet model was used to analyze the echocardiograms to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion regions.

[0012] Furthermore, each of the global max pooling layers is used to independently perform a spatial maximum operation on the feature map in the channel dimension during the feature extraction stage, compressing the feature map to a fixed length to capture multi-scale contextual information at different depths, and obtaining the global pooling features of each layer of the encoder and decoder.

[0013] The output feature of the c-th channel in the i-th stage is expressed by the formula:

[0014]

[0015] in, Indicates the first The output of the first stage Each channel is located in The characteristic response value; These are the row and column indices of the feature map, respectively. This indicates the operation of taking the maximum value at all spatial locations within the channel; This represents the output feature of the channel after global max pooling.

[0016] Furthermore, the classification branch inputs a multi-scale global feature vector, formed by concatenating the global pooling feature vectors from each stage, into the fully connected layer to predict at least two class labels for the lesion region, including:

[0017] The global pooling features of each layer of the encoder and decoder are concatenated and spliced ​​to form a multi-scale global feature vector;

[0018] The multi-scale global feature vector is binary classified using two fully connected layers and Softmax, outputting the classification result of thrombosis / non-thrombosis.

[0019] Furthermore, the encoder-decoder architecture adopted by the CardioMassNet model includes:

[0020] The encoder contains four downsampling stages, each consisting of two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function.

[0021] The decoder includes a corresponding upsampling stage, which uses bilinear interpolation combined with 1×1 convolution to restore the feature map size.

[0022] Furthermore, the loss function of the CardioMassNet model adopts a weighted joint loss, which is expressed by the formula:

[0023]

[0024] in, It is the total loss. It is classification loss. It is a segmentation loss. That is, the weight of the classification loss.

[0025] Furthermore, the segmentation task employs a loss based on the Dice coefficient, expressed by the formula:

[0026]

[0027] in, The classification loss is used to predict the segmentation result. The real mask is , To smooth out the terms, avoid having a denominator of 0.

[0028] Furthermore, the classification task employs cross-entropy loss, expressed by the formula:

[0029]

[0030] in, It is classification loss. This refers to the batch size. Indicates the first Each sample in the true category The predicted probability.

[0031] Furthermore, the performance evaluation of the trained CardioMassNet model using the validation set includes:

[0032] The average Dice coefficient and average intersection-union ratio were used to evaluate the segmentation performance of the model;

[0033] The classification performance of the model was evaluated using accuracy, area under the ROC curve, recall, precision, false positive rate, and F1 score.

[0034] Furthermore, the echocardiogram image dataset is acquired from multiple centers and divided into non-overlapping training, validation, and test sets in chronological order. Data augmentation strategies, including image flipping, rotation, and translation, are applied during the training process.

[0035] On the other hand, the present invention also provides a system for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, comprising:

[0036] The data acquisition module is used to acquire raw echocardiogram images, label and preprocess them to obtain an image dataset, and divide the image dataset into a training set and a validation set.

[0037] The model building module is used to construct the CardioMassNet model based on the U-net architecture. The model adopts an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max pooling layers are configured at the output of each feature extraction stage of the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The output of the model is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch is used to convert the feature map output by the decoder into a pixel-level prediction mask to achieve segmentation of intracardiac space-occupying lesion regions. The classification branch is used to input the multi-scale global feature vector formed by concatenating the global pooling feature vectors of each stage into a fully connected layer to achieve prediction of at least two class labels for the lesion region.

[0038] The training module is used to train the CardioMassNet model using the training set, and save the parameters after training to obtain the trained CardioMassNet model.

[0039] The validation module is used to evaluate the performance of the trained CardioMassNet model using a validation set. After the evaluation, the CardioMassNet model is used to analyze the echocardiogram to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion region.

[0040] Compared with existing technologies, the intracardiac space-occupying lesion region segmentation and classification system and method based on ultrasound images proposed in this invention have the following advantages:

[0041] This invention embeds a global max-pooling layer at the output of each stage of the U-Net encoder-decoder architecture. This concatenates multi-scale feature vectors into the classification branch while retaining features from the decoder's final output to reach the segmentation branch, forming a dual-task architecture with a shared backbone and independent outputs. This design reduces computational cost and overfitting risk through parameter reuse and task regularization, and enhances the ability to discriminate complex patterns through cross-level feature fusion.

[0042] In summary, this invention effectively narrows the interpretation differences between different levels of experience, avoids the risks of traditional enhanced examinations, and provides a scalable intelligent auxiliary analysis solution for ICM lesion region segmentation and classification based on ultrasound images. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images provided by the present invention;

[0044] Figure 2 This is a schematic diagram of the CardioMassNet model's segmentation and classification process for ICM.

[0045] Figure 3 Features of the CardioMass-Net model at different stages Figure 2 3D view;

[0046] Figure 4 This is a comparison chart of the receiver operating characteristic curves of the classification performance of three multi-task models;

[0047] Figure 5 This is a schematic diagram of the intracardiac space-occupying lesion region segmentation and classification system based on ultrasound images provided by the present invention. Detailed Implementation

[0048] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.

[0049] Please see Figure 1This embodiment provides a method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, including the following steps:

[0050] Step S101: Obtain raw echocardiogram images, label and preprocess them to obtain an image dataset, and divide the image dataset into a training set and a validation set;

[0051] Step S102: Construct a CardioMassNet model based on the U-net architecture. The model adopts an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max pooling layers are configured at the output of each feature extraction stage of the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The output of the model is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch is used to convert the feature map finally output by the decoder into a pixel-level prediction mask to achieve segmentation of intracardiac space-occupying lesion areas. The classification branch is used to input the multi-scale global feature vector formed by concatenating the global pooling feature vectors of each stage into a fully connected layer to achieve prediction of at least two class labels for the lesion area.

[0052] Step S103: Train the CardioMassNet model using the training set, and save the parameters after training to obtain the trained CardioMassNet model.

[0053] Step S104: Use the validation set to evaluate the performance of the trained CardioMassNet model. After the evaluation, use the CardioMassNet model to analyze the echocardiogram to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion region.

[0054] This embodiment provides a method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images. It optimizes computational resource utilization through a dual-task joint learning framework. A global max-pooling module is horizontally embedded into the classic U-Net encoder-decoder backbone, constructing a parallel structure of shared encoding and independent decoding. This allows the segmentation and classification tasks to deeply share low-level feature representations, avoiding redundant computation and suppressing overfitting through inter-task regularization. This compresses two models that would otherwise run independently into a single network, significantly reducing deployment costs and inference time. Furthermore, a multi-scale feature stitching strategy allows the classification branch to fuse a complete information chain from low-level texture to high-level semantics, enhancing the ability to discriminate complex lesion patterns. This method avoids the risks of contrast agents, radiation exposure, and increased costs associated with contrast-enhanced or cross-modal imaging techniques. Based entirely on conventional ultrasound images, it provides a more accessible intelligent solution for large-scale bedside rapid assessment.

[0055] In a preferred embodiment, the echocardiogram image dataset is acquired from multiple centers and divided into non-overlapping training, validation, and test sets in chronological order. During training, data augmentation strategies including image flipping, rotation, and translation are applied.

[0056] As a specific example, raw echocardiographic images were acquired using a Philips ultrasound system (IE33, iEElite, EPIQ7, EPIQ7C, Philips Ultrasound, Boswell, ., USA) or a GE ultrasound system (Vivid E9, Vivid E95) equipped with a transthoracic ultrasound probe (frequency range 1-5MHz). Image acquisition followed guidelines to preserve standard sections and clearly show the space-occupying lesion and its relationship to surrounding structures from multiple angles. All images were archived in DICOM format.

[0057] The labels for ICM are determined through a multimodal diagnostic approach, specifically including: (1) surgical or biopsy pathological confirmation (all pathological sections are reviewed by a senior cardiologist with more than 5 years of experience); (2) for non-surgical patients, follow-up ultrasound 3 months after initiation of anticoagulation therapy shows complete disappearance of lesions or significant reduction in volume (>50%); (3) the absence of contrast agent perfusion on enhanced CT or ultrasound can serve as supporting evidence of thrombotic lesions.

[0058] Furthermore, during the training phase, a dataset was constructed using labeled ICM images. Two board-certified echocardiologists (SC or XZ, with 9 and 8 years of experience, respectively) independently delineated regions of interest (ROIs) without knowing the final diagnosis as a reference standard for true segmentation. These ROIs included lesion contours and lesion characteristics (binary classification labels).

[0059] If there are disagreements in delineating the ROI, consult a third senior cardiologist (JW, with 25 years of experience) to reach a consensus and use that consensus as a benchmark.

[0060] The ITK-SNAP tool (Version 4.0.1) was used to segment intracardiac lesions at the pixel level, and the dataset was defined using “thrombosis” and “non-thrombosis” classification labels.

[0061] Images used for model training undergo uniform preprocessing: standard image cropping and scaling to 512×512 pixels; data augmentation strategies include: 1) Flipping: horizontal and vertical flipping; 2) Rotation: clockwise rotation of 30° / 45° / 60°, counterclockwise rotation of 30° / 60°; 3) Translation: translation of 50 pixels to the lower right and 50 pixels to the upper left.

[0062] As a preferred embodiment, the encoder-decoder architecture employed by the CardioMassNet model includes:

[0063] The encoder contains four downsampling stages, each consisting of two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function.

[0064] The decoder includes a corresponding upsampling stage, which uses bilinear interpolation combined with 1×1 convolution to restore the feature map size.

[0065] The following is combined Figure 2 The structure of the CardioMassNet model and its segmentation and classification processes are explained.

[0066] like Figure 2 As shown, the CardioMassNet model performs global feature extraction after each encoder and decoder extracts features in the classic U-Net network, in order to achieve joint learning of segmentation and classification tasks.

[0067] The encoder contains four downsampling stages, each of which includes two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function.

[0068] The decoder includes a corresponding upsampling stage, which uses bilinear interpolation combined with 1×1 convolution to restore the feature map size.

[0069] Specifically, the encoder comprises four downsampling stages, each consisting of two 3×3 convolutions, batch normalization (BN), and a ReLU activation function. Differential dropout values ​​(0.3-0.4) are applied at different depths to enhance generalization and prevent overfitting. A stride of 2 in the convolution achieves a 2x downsampling, reducing spatial resolution layer by layer while preserving the number of channels and extracting higher-level semantic features.

[0070] The batch normalization process is as follows:

[0071]

[0072] in and These are the mean and standard deviation of the current batch of data, respectively. and For learnable scaling and offset parameters, To prevent small constants from being divided by zero, Batch Normalization (BN) can stabilize the training process, accelerate convergence, alleviate gradient vanishing, and improve the model's generalization ability.

[0073] The ReLU activation function is used to enhance the nonlinear expressive power of the model.

[0074] The decoder corresponds to the four stages of the encoder. It uses bilinear interpolation combined with 1×1 convolution to adjust the number of channels and concatenates with the features of the corresponding coding layer to fuse shallow boundary and deep semantic information.

[0075] Bilinear interpolation can be expressed by the following formula:

[0076]

[0077] in, Indicates the input feature map; This indicates that bilinear interpolation is used for upsampling, which increases the spatial resolution of the feature map by a factor of 2. This represents a 1×1 convolution operation, used to adjust the number of channels in the upsampled feature map to match the feature dimension of the corresponding encoder layer. This represents the output feature map after upsampling.

[0078] Bilinear interpolation can efficiently smooth upsampled feature maps, flexibly adjust the number of channels, improve feature representation capabilities, and ensure computational efficiency and training stability.

[0079] Skip connections establish cross-layer connections between the encoder and decoder symmetric layers, concatenating shallow features from the encoder with deep features from the decoder. This preserves spatial details and combines them with high-level semantic information, ensuring that low-level features are fully utilized when restoring spatial resolution and improving segmentation accuracy.

[0080] To enhance the network's global semantic modeling capabilities, in each major stage of the encoder and decoder, namely... Figure 2 Phases 1-9 introduce Global Max Pooling (GMP), which compresses feature maps to a fixed length to capture multi-scale contextual information at different depths, improving the ability to distinguish between overall structure and local details, and providing semantic support for classification tasks.

[0081] In a preferred embodiment, each of the global max pooling layers is used to independently perform a spatial maximum operation on the feature map in the channel dimension during the feature extraction stage, compressing the feature map to a fixed length to capture multi-scale contextual information at different depths, and obtaining the global pooling features of each layer of the encoder and decoder.

[0082] The output feature of the c-th channel in the i-th stage is expressed by the formula:

[0083]

[0084] in, Indicates the first The output of the first stage Each channel is located in The characteristic response value; These are the row and column indices of the feature map, respectively. This indicates the operation of taking the maximum value at all spatial locations within the channel; This represents the output feature of the channel after global max pooling.

[0085] Compared to fully connected layers and Global Average Pooling (GAP), GMP does not require a large number of additional parameters and can extract strong signals in key regions by maximizing the values, making it more sensitive to small targets such as cardiac thrombi. GMP features from different stages are then converged into a new classification branch, achieving multi-scale global semantic modeling. Visual attention heatmaps are shown below. Figure 3 As shown. Figure 3 In the case study, Case 1 was a large non-thrombotic case, and Case 2 was a small thrombotic case.

[0086] As a preferred embodiment, the model's output adopts a dual-branch output strategy, wherein:

[0087] The segmentation branch follows the U-Net design, namely: the decoder output is compressed from 32 channels to a single channel through a 1×1 convolutional layer, and then a pixel-level prediction mask is obtained through the Sigmoid activation function to achieve accurate ROI localization.

[0088] The classification branch is a new global classification branch added by CardioMassNet. It is used to concatenate and stitch together the global pooling features of each layer of the encoder and decoder to form a multi-scale global feature vector, which is then passed through two fully connected layers and Softmax to achieve T (thrombosis) / NT (non-thrombosis) binary classification.

[0089] As a preferred embodiment, the multi-scale global feature vector can be represented as:

[0090]

[0091] in, It is a concatenated multi-scale global feature vector. This represents the global pooling feature vector for the corresponding channel.

[0092] Based on this structure, segmentation and classification tasks share encoded features. On the one hand, this can make full use of information and reduce redundant computation. On the other hand, segmentation and classification learn interactively, with segmentation tasks assisting classification and classification tasks feeding back into segmentation. The two share some parameters, reducing the risk of overfitting and realizing multi-task joint learning of segmentation and classification.

[0093] In a preferred embodiment, the loss function of the CardioMassNet model adopts a weighted joint loss, which is expressed by the formula:

[0094]

[0095] in, It is the total loss. It is classification loss. It is a segmentation loss. That is, the weight of the classification loss.

[0096] This weighted joint loss design promotes feature sharing and collaborative learning, avoids sacrificing the accuracy of one task for optimizing another, effectively suppresses overfitting, and is suitable for multi-task requirements.

[0097] In a preferred embodiment, the segmentation task employs a loss based on the Dice coefficient, expressed by the formula:

[0098]

[0099] in, The classification loss is used to predict the segmentation result. The real mask is , To smooth out the terms, avoid having a denominator of 0.

[0100] In a preferred embodiment, the classification task employs cross-entropy loss, expressed by the formula:

[0101]

[0102] in, It is classification loss. This refers to the batch size. Indicates the first Each sample in the true category The predicted probability.

[0103] In a preferred embodiment, the performance evaluation of the trained CardioMassNet model using a validation set includes:

[0104] The average Dice coefficient and average intersection-union ratio were used to evaluate the segmentation performance of the model;

[0105] The classification performance of the model was evaluated using accuracy, area under the ROC curve, recall, precision, false positive rate, and F1 score.

[0106] As a specific implementation, the model segmentation performance is evaluated using the average Dice coefficient (mDSC) and the average intersection-union ratio (mIoU), expressed by the formula:

[0107]

[0108]

[0109] In the formula, The set of segmentation masks predicted by the model. The set of segmentation masks that are actually labeled.

[0110] The model classification performance was evaluated using accuracy (ACC), area under the ROC curve (AUC), recall (REC), precision (PRE), false positive rate (FPR), and F1 score (F1), expressed by the following formula:

[0111]

[0112]

[0113]

[0114]

[0115]

[0116] In the formula, TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively.

[0117] Furthermore, the preprocessed images and annotations are packaged into a standardized input format for model training. In terms of model architecture, an encoder-decoder framework with U-Net as the backbone is adopted to complete the ICM image segmentation task, effectively capturing feature and detail information and achieving accurate segmentation in ultrasound medical imaging. After iterative training and optimization, the model performance is evaluated based on the loss function and the accuracy on the training set. Finally, based on its performance metrics on the validation set, especially the validation loss, the model weights that achieve the lowest loss on the internal validation set are selected for final evaluation to maximize generality and prevent overfitting.

[0118] Specifically, the CardioMassNet model was implemented using the PyCharm 2025.2.0.1 framework, trained for 150 epochs with the AdamW optimizer (initial learning rate = 1e-4, weight decay = 1e-5) and a batch size of 12. A linear learning rate warm-up strategy was employed, and a cosine annealing scheduler was used after the validation loss plateaued. All experiments were performed on an NVIDIA GeForce RTX 4080 Super (16 GB) GPU, using mixed-precision training to accelerate computation.

[0119] To verify the relative advantages of the CardioMassNet model, two representative multi-task networks were selected as baselines for comparison under completely consistent training, validation, and external test sets, as well as unified image preprocessing, data augmentation, and optimization strategies. MTANet: A multi-task attention network that achieves joint optimization of segmentation and classification by sharing an encoder and introducing an attention mechanism in the task branch, fully utilizing multi-scale features to enhance the representation of task-related regions. Mask R-CNN: A two-stage instance segmentation framework based on a region proposal network, capable of simultaneously generating candidate bounding boxes and pixel-level masks. This study adds a classification branch to the standard structure for thrombosis / non-thrombosis binary classification. All three models use the same loss function (Dice + cross-entropy weighted combination, with consistent weights λ) to ensure a fair comparison.

[0120] Key results include: (a) consistency between the algorithm's segmentation of ICMs and the ROI delineation by sonographers; and (b) the model's ability to classify ICMs into thrombotic and non-thrombotic lesions.

[0121] The model's segmentation performance was assessed using the mean Dice coefficient (mDSC). A two-sample t-test was used to compare the mDSC differences between physicians and the model. Sensitivity and specificity were used to assess the model's classification performance. All statistical analyses were performed using Python 3.10.12. Categorical variables were expressed as percentages and compared using Fisher's exact test or chi-square test according to expected frequencies. Continuous variables were recorded as mean or median. A statistically significant difference was defined as P < 0.05.

[0122] We compared the segmentation and classification performance of CardioMassNet, MTANet, and Mask R-CNN on the internal validation set, internal random test set, and external test set (as shown in Table 1). Its segmentation performance remained the highest across all datasets: internal validation set mDSC 0.779 ± 0.192, internal test set 0.750 ± 0.214, and external test set 0.745 ± 0.172, all higher than MTANet and Mask R-CNN. In classification, CardioMassNet also achieved the best overall performance metrics: internal validation set AUC 0.952 (95% CI 0.927–0.978), internal test set AUC 0.929 (0.885–0.965), and external test set AUC 0.956 (0.890–0.997), while also leading in accuracy and F1 score.

[0123] Table 1. Evaluation of Classification and Segmentation Performance of Different Multi-Task Models

[0124]

[0125] Note: mDSC = mean Dice similarity coefficient; mIOU = mean crossover ratio; SD = standard deviation; ACC = accuracy; AUC = area under the curve; CI = confidence interval; SEN = sensitivity / sensitivity; SPE = specificity; PPV = positive likelihood ratio; NPV = negative likelihood ratio.

[0126] Please see Figure 4 , Figure 4 The ROC curves visually illustrate the classification performance of the three models on different datasets. CardioMassNet demonstrates higher robustness and cross-center generalization ability on all three datasets. As can be seen from the graph, MTANet is inferior in both segmentation and classification, while Mask R-CNN, although close in segmentation, has significantly lower classification performance than CardioMassNet.

[0127] Example 2

[0128] like Figure 5 As shown, this embodiment also provides a system 500 for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, including:

[0129] The data acquisition module 501 is used to acquire raw echocardiogram images, perform annotation and preprocessing to obtain an image dataset, and divide the image dataset into a training set and a validation set.

[0130] The model building module 502 is used to build a CardioMassNet model based on the U-net architecture. The model adopts an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max pooling layers are configured at the output of each feature extraction stage of the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The output of the model is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch is used to convert the feature map finally output by the decoder into a pixel-level prediction mask to achieve segmentation of intracardiac space-occupying lesion areas. The classification branch is used to input the multi-scale global feature vector formed by concatenating the global pooling feature vectors of each stage into a fully connected layer to achieve prediction of at least two class labels for the lesion area.

[0131] Training module 503 is used to train the CardioMassNet model using the training set, and save the parameters after training to obtain the trained CardioMassNet model.

[0132] The validation module 504 is used to evaluate the performance of the trained CardioMassNet model using the validation set. After the evaluation, the CardioMassNet model is used to analyze the echocardiogram to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion region.

[0133] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, characterized in that, include: The raw echocardiogram images were acquired, labeled, and preprocessed to obtain an image dataset, which was then divided into a training set and a validation set. The CardioMassNet model is constructed based on the U-net architecture. The model employs an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max-pooling layers are configured at the output of each feature extraction stage in the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The model's output is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch converts the feature map output by the decoder into a pixel-level prediction mask to segment intracardiac lesion regions. The classification branch inputs the multi-scale global feature vector formed by concatenating the global pooling feature vectors from each stage into a fully connected layer to predict at least two class labels for the lesion region. The CardioMassNet model is trained using the training set, and the parameters are saved after training to obtain the trained CardioMassNet model. The performance of the trained CardioMassNet model was evaluated using a validation set. After the evaluation, the CardioMassNet model was used to analyze the echocardiograms to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion regions.

2. The method according to claim 1, characterized in that, The encoder-decoder architecture used in the CardioMassNet model includes: The encoder contains four downsampling stages, each stage including two 3×3 convolutional layers, a batch normalization layer, and a ReLU activation function; The decoder includes a corresponding upsampling stage, which uses bilinear interpolation combined with 1×1 convolution to restore the feature map size.

3. The method according to claim 2, characterized in that, Each of the global max pooling layers is used to independently perform spatial maximum operation on the feature map in the channel dimension during the feature extraction stage, compressing the feature map to a fixed length to capture multi-scale contextual information at different depths, and obtaining the global pooling features of each layer of the encoder and decoder. The output feature of the c-th channel in the i-th stage is expressed by the formula: ; in, Indicates the first The output of the first stage Each channel is located in The characteristic response value; These are the row and column indices of the feature map, respectively. This indicates the operation of taking the maximum value at all spatial locations within the channel; This represents the output feature of the channel after global max pooling.

4. The method according to claim 2, characterized in that, The classification branch inputs a multi-scale global feature vector, formed by concatenating the global pooling feature vectors from each stage, into the fully connected layer to predict at least two class labels for the lesion region, including: The global pooling features of each layer of the encoder and decoder are concatenated and spliced ​​to form a multi-scale global feature vector; The multi-scale global feature vector is binary classified using two fully connected layers and Softmax, outputting the classification result of thrombosis / non-thrombosis.

5. The method according to claim 1, characterized in that, The loss function of the CardioMassNet model uses a weighted joint loss, expressed by the formula: ; in, It is the total loss. It is classification loss. It is a segmentation loss. That is, the weight of the classification loss.

6. The method according to claim 5, characterized in that, The segmentation task employs a loss based on the Dice coefficient, expressed by the formula: ; in, The classification loss is used to predict the segmentation result. The real mask is , To smooth out the terms, avoid having a denominator of 0.

7. The method according to claim 5, characterized in that, The classification task employs cross-entropy loss, expressed by the formula: ; in, It is classification loss. This refers to the batch size. Indicates the first Each sample in the true category The predicted probability.

8. The method according to claim 5, characterized in that, The performance evaluation of the trained CardioMassNet model using the validation set includes: The average Dice coefficient and average intersection-union ratio were used to evaluate the segmentation performance of the model; The classification performance of the model was evaluated using accuracy, area under the ROC curve, recall, precision, false positive rate, and F1 score.

9. The method according to claim 1, characterized in that, The echocardiogram image dataset was acquired from multiple centers and divided into non-overlapping training, validation, and test sets in chronological order. Data augmentation strategies, including image flipping, rotation, and translation, were applied during the training process.

10. A system for segmenting and classifying intracardiac space-occupying lesions based on ultrasound images, characterized in that, include: The data acquisition module is used to acquire raw echocardiogram images, label and preprocess them to obtain an image dataset, and divide the image dataset into a training set and a validation set. The model building module is used to construct the CardioMassNet model based on the U-net architecture. The model adopts an encoder-decoder architecture, with each encoder and decoder containing multiple cascaded feature extraction stages. Global max pooling layers are configured at the output of each feature extraction stage of the encoder and decoder to capture contextual information at different scales and generate corresponding global feature vectors. The output of the model is configured as a dual-branch output module including a segmentation branch and a classification branch. The segmentation branch is used to convert the feature map output by the decoder into a pixel-level prediction mask to achieve segmentation of intracardiac space-occupying lesion regions. The classification branch is used to input the multi-scale global feature vector formed by concatenating the global pooling feature vectors of each stage into a fully connected layer to achieve prediction of at least two class labels for the lesion region. The training module is used to train the CardioMassNet model using the training set, and save the parameters after training to obtain the trained CardioMassNet model. The validation module is used to evaluate the performance of the trained CardioMassNet model using a validation set. After the evaluation, the CardioMassNet model is used to analyze the echocardiogram to be detected and output the segmentation and classification results of the intracardiac space-occupying lesion region.