Diffuse large B-cell lymphoma segmentation system based on deep learning

By developing a deep learning-based segmentation system for diffuse large B-cell lymphoma, this system utilizes a bi-branch encoder, cross-modal feature cross-fusion, and a global-inhibition attention module to address the shortcomings of global context modeling and multi-modal feature fusion. This results in accurate segmentation of diffuse large B-cell lymphoma, improving the accuracy and robustness of the segmentation.

CN120876853APending Publication Date: 2025-10-31ZHEJIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510955642.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing deep learning-based segmentation methods for diffuse large B-cell lymphoma suffer from insufficient global context modeling capabilities, inadequate extraction of multimodal differential features and utilization of complementary information, and an insufficient balance between suppressing redundant information and preserving key features. These issues result in segmentation results that are deficient in terms of anatomical rationality and accuracy.

Method used

A deep learning-based segmentation system for diffuse large B-cell lymphoma is employed, comprising a data acquisition module, a dual-branch encoder module, a global-inhibition attention module, a cross-modal feature fusion module, and a multi-scale skip connection module. The dual-branch encoder processes PET and CT images separately, the cross-modal feature fusion module enables dynamic complementary fusion of PET and CT features, the global-inhibition attention module dynamically filters and suppresses redundant information, and the multi-scale skip connection module restores spatial details to achieve accurate segmentation.

Benefits of technology

The segmentation method significantly improved the accuracy and robustness of diffuse large B-cell lymphoma segmentation, with a Dice similarity coefficient of 0.892, a 95% Hausdorff distance of 27.560, a mean symmetry surface distance of 5.554, and recall and precision of 0.908 and 0.914, respectively. This demonstrates high sensitivity and anti-interference ability for key lesions, and the segmentation results have good potential for clinical application.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876853A_ABST
    Figure CN120876853A_ABST
Patent Text Reader

Abstract

The invention discloses a diffuse large B-cell lymphoma segmentation system based on deep learning, and the system comprises a data collection module which is used for collecting PET and CT images of a patient; the double-branch encoder module is used for processing PET and CT images respectively; the cross-modal feature cross fusion module is used for extracting positioning clues from the CT anatomical features by taking the PET metabolic features as queries; taking CT structure characteristics as query, and screening PET high-metabolism regions; the global-attention suppression module is used for extracting global context features, performing down-sampling, remodeling input features into a low-dimensional matrix, then performing self-attention calculation, generating a sparse weight matrix for the input features based on a full connection layer, and then performing sparse attention calculation; and finally, fusing the two attention calculation results and outputting enhanced features. And the multi-scale jump connection module is used for fusing the encoder features and the decoder features and recovering spatial details.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image segmentation technology, and in particular to a deep learning-based segmentation system for diffuse large B-cell lymphoma. Background Technology

[0002] Diffuse large B-cell lymphoma (DLBCL), the most common subtype of non-Hodgkin's lymphoma, requires precise segmentation for treatment planning and prognostic assessment. In recent years, PET / CT-based multimodal imaging analysis has become a core tool for the diagnosis and staging of DLBCL. PET reflects tumor functional characteristics through metabolic activity information (such as FDG uptake), while CT provides high-resolution anatomical information. The fusion of these two technologies significantly improves the ability to locate lesions and determine their boundaries, further driving the development of automated segmentation techniques.

[0003] At the technical level, deep learning has become the mainstream paradigm for medical image segmentation.

[0004] 1. The encoder-decoder architecture represented by U-Net (Ronneberger O, Fischer P, Brox TU-net: Convolutional networks for biomedical image segmentation[C] / / Medical image computing and computer-assisted intervention-MICCAI 2015:18th international conference,Munich,Germany,October 5-9,2015,proceedings,part III 18.Springer internationalpublishing,2015:234-241.) achieves multi-scale feature fusion through skip connections, demonstrating significant advantages in lymphoma segmentation tasks.For example, Blanc-Durand et al. (Blanc-Durand P, Jégou S, Kanoun S, et al. Fully automatic segmentation of diffuse large B cell lymphoma lesions on 3D FDG-PET / CT for total metabolic tumor volume prediction using a convolutional neural network[J]. European Journal of Nuclear Medicine and Molecular Imaging, 2021, 48: 1362-1370.) achieved fully automatic segmentation of DLBCL based on 3D U-Net, and its TMTV prediction results were highly consistent with clinical assessments; Yousefirizi et al. (Yousefirizi F, Klyuzhin I S, OJH, et al. TMTV-Net: fully automated total metabolic tumor volumesegmentation in lymphoma PET / CT images—a multi-center generalizability analysis[J]. European Journal of Nuclear Medicine and Molecular Imaging, 2024, 51(7): 1937-1954.) achieved fully automatic segmentation of DLBCL based on cascaded multi-resolution 3D U-Net. U-Net and semi-supervised strategies effectively alleviate the problem of insufficient labeled data.

[0005] 2. Furthermore, regarding the fusion methods for multimodal data, scholars have proposed three strategies: pre-fusion (input-level fusion), mid-fusion (feature-level fusion), and post-fusion (decision-level fusion). Input-level fusion involves directly stitching PET and CT images into a multi-channel input and extracting joint features through a shared encoder. This method is computationally efficient but may ignore the heterogeneity between modalities. Feature-level fusion involves using a dual-branch encoder to process PET and CT features independently, and then interacting through cross-attention (such as MMCA-Net (Zhao W, Huang Z, Tang S, et al.Mmca-net: a multimodal cross attentiontransformer network for nasopharyngeal carcinoma tumor segmentation based on a total-body PET / CT system[J].IEEE Journal of Biomedical and HealthInformatics,2024.)) or feature map stitching. This type of method can preserve modality-specific information but requires the design of complex fusion modules. Decision-level fusion involves training single-modal segmentation networks for PET and CT separately, and generating the final result through weighted averaging or confidence fusion. Its advantage lies in its high flexibility, but it may lose fine-grained spatial information. Huang et al. (Huang L, T, Tonnelet D, et al. Deep PET / CT fusion with Dempster-Shafertheory for lymphoma segmentation[C] / / Machine Learning in Medical Imaging:12th International Workshop,MLMI 2021,Held in Conjunction with MICCAI 2021,Strasbourg,France,September 27,2021,Proceedings 12.Springer InternationalPublishing,2021:30-39.) By integrating the single-modal segmentation results of PET and CT through evidence fusion layer, the Dice coefficient was improved by 3.2%; Yuan et al. (Yuan C,Zhang M,Huang X, et al.Diffuse large B-cell lymphomasegmentation in PET-CT images via hybrid learning for feature fusion[J].Medical Physics, 2021, 48(7):3665-3678.) Design a dual encoder structure, combined with spatial feature cross-connection, to fully explore the complementarity of PET and CT.

[0006] Semi-supervised and unsupervised learning techniques have also become research hotspots. Li et al. (Li H, Jiang H, Li S, et al. DenseX-net: an end-to-end model for lymphoma segmentation in whole-body PET / CT images[J]. IEEE Access, 2019, 8: 8004-8018.) proposed a hybrid learning framework based on reconstruction flow, which utilizes unlabeled data to enhance the model's generalization ability; Yousefirizi (Yousefirizi F, Shiri I, OJH, et al. Semi-supervised learning towards automated segmentation of PET images with limited annotations: application to lymphoma patients[J]. Physical and Engineering Sciences in Medicine, 2024, 47(3): 833-849.) introduced fuzzy clustering and boundary loss functions to achieve robust segmentation of DLBCL in small sample scenarios. These works show that by reducing the dependence on fully labeled data, the applicability of the model in clinical practice can be significantly improved.

[0007] Existing deep learning-based medical image segmentation methods still have the following drawbacks:

[0008] (1) Insufficient global context modeling capability:

[0009] Current deep learning segmentation methods for diffuse large B-cell lymphoma (DLBCL) are mostly based on traditional U-Net and its variants. Traditional U-Net relies entirely on convolutional neural networks (CNNs) for feature extraction. Due to the inherent characteristics of their local receptive fields, CNNs struggle to effectively capture the systemic distribution characteristics of DLBCL lesions. DLBCL lesions often exhibit extensive involvement (e.g., bone marrow, spleen, and extranodal tissues) and high morphological heterogeneity. Traditional U-Nets and their variants, such as ResNet++ and AttUNet, extract features only through local convolutional operations, failing to establish long-range spatial dependencies. In a comparative experiment using 152 newly diagnosed and pathologically confirmed DLBCL patients from the Second Affiliated Hospital of Zhejiang University School of Medicine between 2013 and 2019, the ResNet++ model based on a pure CNN architecture achieved a DSC of only 0.778 and an HD95 as high as 52.75 in the segmentation of scattered small lesions, indicating significant deficiencies in global localization and boundary coherence. In addition, traditional methods are insufficient in modeling the spatial correlation of lesions throughout the body, which leads to deviations in the anatomical rationality of the segmentation results and the defect of mistakenly including adjacent normal tissues in the lesion area.

[0010] (2) Insufficient extraction of multimodal differential features and utilization of complementary information:

[0011] Existing PET / CT fusion strategies mostly employ simple stitching or static weighted fusion, failing to dynamically balance the differential contributions of metabolic activity in PET images and anatomical structures in CT images. Specifically:

[0012] 1) In the multi-channel input stage, PET and CT images are directly stitched together while ignoring the heterogeneity between modalities, which leads to feature confusion problems. For example, the spatial misalignment between high metabolic regions in PET and low density regions in CT affects the segmentation accuracy.

[0013] 2) Although some models extract features from different modalities through dual-branch encoders, they lack cross-modal interaction mechanisms and cannot achieve dynamic feature complementarity in key areas such as lesions with blurred boundaries. Experiments show that the current best method, nnUNet, has a DSC of 0.885 when fusing PET / CT features, but still suffers from missegmentation due to insufficient utilization of modal complementarity (VOE = 0.206).

[0014] (3) Insufficient balance between suppressing redundant information and preserving key features:

[0015] Ordinary attention mechanisms introduce a lot of background noise into global computation and lack the ability to adaptively suppress redundant information.

[0016] Specifically, this manifests as follows:

[0017] 1) Dense attention mechanisms capture long-range dependencies through global computation, but do not dynamically suppress low-weight features, resulting in irrelevant noise or background information affecting computational efficiency and accuracy.

[0018] 2) In sparse attention mechanisms, conventional static threshold filtering cannot adjust the information retention ratio according to the characteristics of the input data, which can easily lead to the over-removal of key features. Summary of the Invention

[0019] The purpose of this invention is to address the shortcomings of existing technologies by proposing a deep learning-based segmentation system for diffuse large B-cell lymphoma.

[0020] The objective of this invention is achieved through the following technical solution: a deep learning-based system for segmenting diffuse large B-cell lymphoma, comprising a data acquisition module, a dual-branch encoder module, a global-inhibition attention module, a cross-modal feature cross-fusion module, and a multi-scale skip connection module;

[0021] The data acquisition module is used to acquire PET and CT images of patients diagnosed with diffuse large B-cell lymphoma and to construct a dataset;

[0022] The dual-branch encoder module is used to process PET and CT images separately and perform feature encoding.

[0023] The cross-modal feature fusion module is used to realize bidirectional interaction between PET and CT features in the dual-branch encoder module. It uses PET metabolic features as a query to extract localization clues from CT anatomical features and uses CT structural features as a query to filter PET high metabolic regions.

[0024] The global-inhibition attention module is used to extract global context features processed by the cross-modal feature cross-fusion module and downsample them; specifically, the input features are reshaped into a low-dimensional matrix, and then self-attention calculation is performed; for the input features, a sparse weight matrix is ​​generated based on the fully connected layer, and then sparse attention calculation is performed; finally, the two attention calculation results are fused and the enhanced features are output.

[0025] The multi-scale skip connection module is used to fuse encoder features and decoder features to restore spatial details and gradually convert the encoded features into segmentation results of lymphoma lesions. That is, it outputs predictions of the location, shape and extent of lymphoma lesions in the image, thereby achieving accurate segmentation of diffuse large B-cell lymphoma.

[0026] Furthermore, the dual-branch encoder module has a hierarchical structure, with each branch containing multiple stacked global-suppression attention modules.

[0027] Furthermore, the cross-modal feature fusion module receives input from the encoder's PET and CT feature maps. Based on a bidirectional cross-attention mechanism, it achieves dynamic complementary fusion of PET metabolic activity and CT anatomical structure, specifically as follows:

[0028] First, within an attention branch, attention weights are calculated using PET embeddings as queries and CT embeddings as keys and values.

[0029] Meanwhile, in another attention branch, attention weights are calculated using CT embeddings as queries and PET embeddings as keys and values.

[0030] Finally, the features from the two paths are fused to obtain the cross-fused features; and residual connections are used to merge the original input of the PET path with the cross-fused features to obtain the final fused features.

[0031] Furthermore, the global-suppression attention module includes a global information aggregation branch and a key feature filtering branch; the global information aggregation branch is used to capture global context information and perform self-attention calculation, and the key feature filtering branch dynamically suppresses redundant information and performs sparse attention calculation through sparse weights; the fusion ratio of the two branches is controlled by learnable parameters, and the spatial and channel attention outputs are concatenated, and the feature map resolution is restored through convolution operations.

[0032] Furthermore, the global information aggregation branch is specifically implemented as follows:

[0033] (1) Input feature mapping: Input features Reshaped into a low-dimensional matrix Where m = H × W × D

[0034] (2) Self-attention calculation

[0035] Spatial attention:

[0036]

[0037] Channel attention:

[0038]

[0039]

[0040] Furthermore, the key feature filtering branch is specifically implemented as follows:

[0041] (1) Input features X are processed by a fully connected layer (FC) to generate a sparse weight matrix:

[0042] W sparse =Softmax(FC(X))

[0043] Spatial sparse weights With channel sparse weights It acts on the spatial and channel dimensions respectively;

[0044] (2) Sparse attention calculation:

[0045] Space filtering:

[0046]

[0047] Channel filtering:

[0048]

[0049] The beneficial effects of this invention are:

[0050] 1. This invention is based on a cross-modal feature fusion mechanism, which breaks through the limitations of traditional single-modal dominance or simple feature superposition. It realizes dynamic complementary fusion of PET metabolic activity and CT anatomical structure, and solves the problem of feature confusion caused by the heterogeneity of PET and CT images in terms of physical expression such as resolution and noise distribution, as well as semantic expression such as metabolic function and anatomical morphology. At the same time, it solves the problem that the modality-dominated fusion strategy cannot dynamically balance the contributions of multiple modalities, resulting in significant segmentation errors in fuzzy regions.

[0051] 2. This invention is based on the dual-branch collaborative mechanism of "global information aggregation" and "key feature selection" of the global-suppression attention module. Through dynamic sparse weight mechanism and multi-dimensional feature selection, it solves the shortcomings of traditional attention mechanism in suppressing redundant information and retaining key features.

[0052] 3. This invention optimizes global-inhibitory attention simultaneously in the spatial dimension (locating lesion areas) and the channel dimension (screening discriminative features) through multi-dimensional feature optimization. It optimizes feature expression from the dual perspectives of anatomical structure and metabolic activity, further improving the segmentation robustness of small lesions and fuzzy boundaries, and significantly enhancing the sensitivity and anti-interference ability of key lesions in the DLBCL segmentation task. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 This is a structural diagram of the overall model of the present invention;

[0055] Figure 2This is a structural diagram of the PET / CT feature cross-fusion module of the present invention;

[0056] Figure 3 This is a structural diagram of the global-inhibition attention module of the present invention;

[0057] Figure 4 This is a flowchart of the global-suppression attention module calculation process of the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described below with reference to the accompanying drawings and examples. It should be understood that the specific examples described herein are merely illustrative and not intended to limit the invention.

[0059] This invention proposes a deep learning-based segmentation system for diffuse large B-cell lymphoma, aiming to address the core shortcomings of traditional techniques in global context modeling, dynamic fusion of multimodal features, and suppression of redundant information. The following is illustrated in the accompanying figures. Figures 1 to 4 The details of the process and its technical aspects will be elaborated upon.

[0060] like Figure 1 As shown, the core architecture of this invention is a dual-branch encoder-decoder network, which includes the following key modules:

[0061] Dual-branch encoder module: processes PET and CT images separately. Each branch consists of a multi-level feature extraction block (GSA module) and a downsampling module.

[0062] Cross-modal feature fusion module (CPCFM): Enables bidirectional dynamic interaction between PET and CT features at each level of the encoder.

[0063] Global-Suppressive Attention Module (GSA): Optimizes feature selection in the encoder and decoder, suppresses redundant information, and preserves key lesion features.

[0064] Multi-scale jump connection module: fuses encoder features with decoder features to restore spatial details.

[0065] Working in conjunction with the encoder, the decoder receives features from the dual-branch encoder module after processing PET and CT images for diffuse large B-cell lymphoma segmentation. These features contain key information such as metabolic activity and anatomical structure. The decoder, in cooperation with the multi-scale skip connection module, fuses the encoder features with its own features to restore spatial details, gradually converting the encoded features into segmentation results for lymphoma lesions. This output predicts the location, shape, and extent of lymphoma lesions in the image, thus achieving accurate segmentation of diffuse large B-cell lymphoma.

[0066] The specific implementations of each of the above modules are as follows:

[0067] 1. Dual-branch encoder module

[0068] (1) Hierarchical structure

[0069] Each branch contains N1 stages, and each stage consists of the following components.

[0070] 1) Global and Suppression Attention (GSA) Module Stacking: Each stage contains N1 stacked Global and Suppression Attention (GSA) modules, which extract global context through self-attention mechanism and dynamically suppress redundant information.

[0071] 2) Downsampling layer: The feature map resolution is reduced to half of its original value through stride convolution.

[0072] (2) Cross-modal fusion

[0073] After each encoder output, a cross-modal feature fusion module (CPCFM) enables bidirectional interaction between PET and CT features. Using PET metabolic features as a query, localization clues are extracted from CT anatomical features; using CT structural features as a query, high-metabolic regions in PET are filtered, thereby improving the ability to extract differentiated information between different modalities.

[0074] 2. Cross-modal feature fusion module (CPCFM)

[0075] like Figure 2As shown, the CPCFM module overcomes the limitations of traditional single-modality-dominated or simple feature superposition through a bidirectional cross-attention mechanism. By constructing a bidirectional query-key-value interaction framework between PET and CT features, it achieves dynamic complementary fusion of PET metabolic activity and CT anatomical structure. This solves the problem of feature confusion caused by the heterogeneity of PET and CT images in terms of physical representation (resolution, noise distribution, etc.) and semantic representation (metabolic function, anatomical morphology, etc.) in traditional methods. It also addresses the issue that modality-dominated fusion strategies cannot dynamically balance multimodal contributions, leading to significant segmentation errors in ambiguous regions. Unlike unidirectional auxiliary or early / late fusion strategies, the PET / CT feature cross-fusion (CPCF) module allows PET and CT to be query subjects for each other. For example, when PET is the query, CT anatomical features provide spatial localization support as key-value pairs; conversely, when CT is the query, PET metabolic features assist in the discrimination of functionally active regions. This bidirectional interaction mechanism can dynamically capture the differences and complementarities between different modalities, significantly enhancing the resolution of lesion boundaries. Through cross-modal attention computation, the PET / CT feature cross-fusion (CPCF) module can adaptively adjust the fusion weights based on the local characteristics of the lesion (such as metabolic heterogeneity or anatomical ambiguity), thereby achieving more accurate feature alignment in complex scenarios (such as small lesions or organ infiltration) and effectively suppressing background interference. The specific calculation process is as follows:

[0076] (1) Feature embedding:

[0077] The input features are PET feature maps from the encoder. and CT feature map Where H, W, and D are spatial dimensions, and C is the number of channels.

[0078] The module extracts deep features through independent convolutional layers.

[0079] (2) Bidirectional cross-attention mechanism

[0080] First, using PET embedding as the query (QPET = EPETWQPET), CT embedding as the key and value (KCT = ECTWKCT, VCT = ECTWVCT), we calculate the attention weights:

[0081]

[0082] Meanwhile, in another attention branch, attention weights are calculated using CT embedding as the query (QCT = ECTWQCT), PET embedding as the key and value (KPET = EPETWKPET), and VPET = EPETWVPET).

[0083]

[0084] (3) Feature fusion and residual connection

[0085] Cross-fusion of features from two paths

[0086] F concat =Concat(Attn) PET→CT Attn CT→PET (3)

[0087] Finally, to obtain richer feature representations, this invention employs residual connections to merge the original input of the PET path with cross-fused features, resulting in the final fused features.

[0088] Output = Conv 3d (F concat ,F PET (4)

[0089] 3. Global-Suppressive Attention Module (GSA)

[0090] like Figure 3 As shown, Global and Suppression Attention (GSA) is the core module of this invention. It aims to address the shortcomings of traditional attention mechanisms in suppressing redundant information and preserving key features through a dynamic sparse weighting mechanism and multi-dimensional feature selection, based on a dual-branch collaborative mechanism of "global information aggregation" and "key feature selection." The Global Information Aggregation (GIA) branch comprehensively captures long-range contextual dependencies through a standard self-attention mechanism, ensuring that potentially important features (such as scattered lesions or extensive organ involvement) are not missed. The Key Feature Selection (KFS) branch introduces a sparse weighting mechanism, dynamically generating a probability distribution based on input features to suppress low-confidence regions (such as noise or normal tissue) and strengthen the attention weight of high-confidence lesion regions. Adaptive weight balancing dynamically adjusts the contribution ratio of GIA and KFS through learnable parameters, enabling the model to autonomously choose to emphasize global perception or local focus based on data characteristics (such as lesion distribution sparsity or metabolic heterogeneity), avoiding excessive suppression of effective information. Its core design includes the following three parts:

[0091] (1) Global Information Aggregation Branch (GIA)

[0092] Function: Capture global contextual information to avoid missing key features due to limitations in local receptive fields;

[0093] Implementation method:

[0094] ① Input feature mapping: Input features Reshaped into a low-dimensional matrix Where m = H × W × D

[0095] ② Self-attention calculation

[0096] Spatial attention:

[0097]

[0098]

[0099] Channel attention:

[0100]

[0101] (2) Key Feature Filtering Branch (KFS)

[0102] Function: Dynamically suppresses redundant information through sparse weights to focus on key lesion areas.

[0103] Implementation method:

[0104] Sparse weight generation:

[0105] The input feature X is passed through a fully connected layer (FC) to generate a sparse weight matrix:

[0106] W sparse =Softmax(FC(X)) (13)

[0107] Spatial sparse weights With channel sparse weights They act on the spatial and channel dimensions respectively.

[0108] Sparse attention computation:

[0109] Space filtering:

[0110]

[0111] Channel filtering:

[0112]

[0113] (3) Adaptive weight fusion

[0114] Function: Dynamically balances the contribution of global information and key feature selection.

[0115] Implementation method:

[0116] ① Weight parameterization: The fusion ratio of the global branch (GIA) and the selection branch (KFS) is controlled by learnable parameters α and β.

[0117] Attn final-spatial =α·Attn GIA-spatial +β·Attn KFS-spatial (16)

[0118] Attn final-channel =α·Attn GIA-channel +β·Attn KFS-channel (17)

[0119] ② Multi-dimensional feature fusion: Spatial and channel attention outputs are concatenated, and feature map resolution is restored through convolution operations.

[0120]

[0121] The GSA module calculation flowchart is as follows: Figure 4 The calculation process can be summarized as follows:

[0122] 1) Input PET / CT feature map X

[0123] 2) Relying on the GIA branch, the global context is calculated through self-attention.

[0124] 3) Relying on KFS branches, sparse weights are generated and background noise is dynamically suppressed.

[0125] 4) The two types of attention results are fused using learnable parameters α and β.

[0126] 5) The convolution operation restores the resolution and outputs the enhanced feature map F. out

[0127] The GSA module achieves multi-dimensional feature optimization through a dual-branch design of global information aggregation and dynamic sparse filtering. It simultaneously implements global-inhibitory attention in both the spatial dimension (locating lesion regions) and the channel dimension (filtering discriminative features), optimizing feature expression from both anatomical structure and metabolic activity perspectives. This further enhances the segmentation robustness of small lesions and ambiguous boundaries, significantly improving the sensitivity and anti-interference ability for key lesions in the DLBCL segmentation task. Its adaptive weight fusion mechanism provides a universal feature enhancement paradigm for medical image analysis.

[0128] The deep learning-based segmentation system for diffuse large B-cell lymphoma provided by this invention demonstrates significant advantages in several key evaluation metrics. Specifically, its Dice similarity coefficient (DSC) reaches 0.892, indicating a high degree of overlap between the segmented result and the actual lesion area; the 95% Hausdorff distance (HD95) is 27.560, and the average symmetric surface distance (ASSD) is 5.554, reflecting good boundary segmentation accuracy; the recall rate is 0.908 and the precision is 0.914, showing efficient capture of actual lesions and high prediction accuracy; the voxel overlap error (VOE) is 0.169 and the relative volume difference (RVD) is 0.155, indicating small segmentation errors. This system can accurately segment diffuse large B-cell lymphoma lesions and has good potential for clinical application.

[0129] A specific embodiment of the present invention is as follows:

[0130] (1) Data set and preprocessing method

[0131] To verify the effectiveness of the proposed method, this invention constructed a dataset of PET and CT images of patients diagnosed with diffuse large B-cell lymphoma (DLBCL) for evaluation. This dataset, sourced from a hospital, includes 152 newly diagnosed and pathologically confirmed nasopharyngeal carcinoma (DLBCL) patients from 2013 to 2019. Inclusion criteria included: 1) pathologically confirmed DLBCL; 2) age over 18 years; 3) prior to treatment... 18 F) Fluorodeoxyglucose (FDG) positron emission tomography / computed tomography (PET / CT); 4) Initial treatment with the R-CHOP regimen (rituximab, cyclophosphamide, doxorubicin, vincristine, and prednisone) or the R-EPOCH regimen (rituximab, etoposide, prednisone, vincristine, cyclophosphamide, and doxorubicin). Exclusion criteria included patients with concomitant central nervous system lymphoma or other malignancies, or those with incomplete follow-up.

[0132] All included patient PET images were analyzed by two experienced nuclear medicine physicians unaware of the patients' prognosis. Volumes of interest (VOIs) were semi-automatically delineated using LIFEx software (version 6.30, https: / / www.lifexsoft.org / index.php) with a fixed 41% SUVmax threshold. To minimize the influence of partial volume effects, lesions with a diameter of at least 2 cm on CT or those not readily apparent on CT but with a metabolic volume of at least 4.2 cm were selected. 3The lesions are analyzed. If focal or multifocal lesions are found in the bone marrow and uptake is higher than in the liver, bone marrow involvement is considered. This allows for the determination of the complete lymphoma lesion area as the gold standard.

[0133] To eliminate the influence of factors such as weight, injection dosage, and scan time, and to achieve quantitative comparisons among different patients and under different scanning conditions, this invention converts all PET scan results into standardized uptake values ​​(SUVs). To match PET and CT images and voxels, this invention uses linear interpolation to resample each pair of registered PET / CT scan images to a size of 200*200*n (n ranging from 326 to 536) at spatial resolution, with a slice thickness of 3.0 mm. Isotropic spacing is used in three-dimensional space. Next, this invention performs data augmentation through scaling, random flipping, and rotation to avoid overfitting and improve training robustness. Finally, all training data are normalized to a distribution with a mean of zero and a standard deviation of one.

[0134] All data were randomly divided into training and testing sets in a 7:3 ratio. The training set included 107 volumetric images, and the testing set included 45 volumetric images. These volumetric images were used for training and testing the 3D model.

[0135] (2) Experimental details

[0136] The experiments were implemented using PyTorch 11.3 and Python 3.8. To ensure fair comparison, all models used the same input size, preprocessing methods, and training loss. These models were trained on a single NVIDIA A6000 GPU with 40GB of VRAM for 1000 epochs, with a learning rate of 0.01 and a weight decay factor of 3 × 10⁻⁶. -5 The batch size is set to 2, and the optimizer used is the Adam optimizer.

[0137] (3) Evaluation indicators

[0138] This invention employs multiple evaluation metrics to comprehensively measure the quality and accuracy of segmentation results. The most important reference metric is the Dice similarity coefficient (DSC), which assesses the overlap between the predicted results and the ground truth annotations; a higher value indicates a higher similarity between the segmented result and the real target. The 95% Hausdorff distance (HD95) measures the distance between the farthest points in two sets of points; a smaller HD95 value indicates more accurate segmentation at boundaries. Recall and precision reflect the segmentation method's ability to capture the real target and the accuracy of the predicted results, respectively. Voxel overlap error (VOE) quantifies the error in the segmentation results; a smaller VOE value represents a more accurate segmentation. Relative volume difference (RVD) assesses the volume difference between the segmented result and the ground truth annotations; a smaller RVD value indicates more precise segmentation. The average symmetric surface distance (ASSD) measures the distance between the segmented surface and the real target surface; a smaller ASSD value indicates more accurate surface segmentation. These metrics help this invention objectively evaluate the success of the segmentation technique.

[0139] This invention achieves a comprehensive performance breakthrough in the segmentation task of diffuse large B-cell lymphoma (DLBCL) through a CNN-Transformer hybrid architecture, dynamic cross-modal fusion (CPCFM), and a global-inhibition attention (GSA) module. Experimental data shown in the table below demonstrate that this invention significantly outperforms existing methods in key metrics such as overlap (DSC), boundary accuracy (HD95, ASSD), sensitivity, and specificity (Recall, Precision), exhibiting clear clinical translational potential and technological innovation.

[0140] DSC↑ HD95↓ Recall↑ Precision↑ VOE↓ RVD↓ ASSD↓ nnUNet 0.885 44.950 0.866 0.913 0.206 0.201 5.930 AttUNet 0.775 68.154 0.851 0.796 0.332 0.351 16.145 TransUNet 0.793 58.799 0.836 0.831 0.310 0.314 11.001 SwinUNET 0.829 48.671 0.878 0.823 0.284 0.290 11.227 ResNet++ 0.778 52.746 0.825 0.799 0.355 0.383 19.5939 nnFormer 0.828 48.890 0.820 0.874 0.259 0.280 19.910 UNETR++ 0.848 64.500 0.875 0.870 0.219 0.285 20.940 This invention 0.892 27.560 0.908 0.914 0.169 0.155 5.554

[0141] The above embodiments are used to explain and illustrate the present invention, but not to limit the present invention. Any modifications and changes made to the present invention within the spirit and scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A deep learning-based segmentation system for diffuse large B-cell lymphoma, characterized in that, The system includes a data acquisition module, a dual-branch encoder module, a global-inhibition attention module, a cross-modal feature cross-fusion module, and a multi-scale skip connection module; The data acquisition module is used to acquire PET and CT images of patients diagnosed with diffuse large B-cell lymphoma and to construct a dataset; The dual-branch encoder module is used to process PET and CT images separately and perform feature encoding. The cross-modal feature fusion module is used to realize bidirectional interaction between PET and CT features in the dual-branch encoder module. It uses PET metabolic features as a query to extract localization clues from CT anatomical features and uses CT structural features as a query to filter PET high metabolic regions. The global-inhibition attention module is used to extract global context features processed by the cross-modal feature cross-fusion module and downsample them; specifically, the input features are reshaped into a low-dimensional matrix, and then self-attention calculation is performed; for the input features, a sparse weight matrix is ​​generated based on the fully connected layer, and then sparse attention calculation is performed; finally, the two attention calculation results are fused and the enhanced features are output. The multi-scale skip connection module is used to fuse encoder features and decoder features to restore spatial details and gradually convert the encoded features into segmentation results of lymphoma lesions. That is, it outputs predictions of the location, shape and extent of lymphoma lesions in the image, thereby achieving accurate segmentation of diffuse large B-cell lymphoma.

2. The deep learning-based segmentation system for diffuse large B-cell lymphoma according to claim 1, characterized in that, The dual-branch encoder module has a hierarchical structure, with each branch containing multiple stacked global-suppression attention modules.

3. The deep learning-based segmentation system for diffuse large B-cell lymphoma according to claim 1, characterized in that, The cross-modal feature fusion module takes input from the encoder's PET and CT feature maps. Based on a bidirectional cross-attention mechanism, it achieves dynamic complementary fusion of PET metabolic activity and CT anatomical structure, as specifically implemented below: First, within an attention branch, attention weights are calculated using PET embeddings as queries and CT embeddings as keys and values. Meanwhile, in another attention branch, attention weights are calculated using CT embeddings as queries and PET embeddings as keys and values. Finally, the features from the two paths are fused to obtain the cross-fusion features; The original input of the PET path is then merged with the cross-fusion features using residual connections to obtain the final fused features.

4. The deep learning-based segmentation system for diffuse large B-cell lymphoma according to claim 1, characterized in that, The global-suppression attention module includes a global information aggregation branch and a key feature filtering branch; The global information aggregation branch is used to capture global context information and perform self-attention calculation. The key feature filtering branch dynamically suppresses redundant information and performs sparse attention calculation through sparse weights. The fusion ratio of the two branches is controlled by learnable parameters, and the spatial and channel attention outputs are concatenated. The feature map resolution is restored through convolution operations.

5. A deep learning-based segmentation system for diffuse large B-cell lymphoma according to claim 4, characterized in that, The global information aggregation branch is implemented as follows: (1) Input feature mapping: Input features Reshaped into a low-dimensional matrix Where m = H × W × D (2) Self-attention calculation Spatial attention: Channel attention:

6. A deep learning-based segmentation system for diffuse large B-cell lymphoma according to claim 4, characterized in that, The key feature filtering branch is implemented as follows: (1) Input features X are processed by a fully connected layer (FC) to generate a sparse weight matrix: W sparse =Softmax(FC(X)) Spatial sparse weights With channel sparse weights It acts on the spatial and channel dimensions respectively; (2) Sparse attention calculation: Space filtering: Channel filtering:

Citation Information

Cited By

  • Lymphoma-oriented FDG PET-CT focus automatic segmentation and scoring method and system

    CN121616614A

  • Bone defect repair material suitability prediction method based on deep learning

    CN121862427A