Pulmonary nodule segmentation method and system based on diffusion model

Through the latent spatial fusion and multimodal FSUNet architecture based on anatomical constraints, combined with the Dice-Focal loss function, the problems of boundary accuracy and small target area segmentation in lung nodule segmentation are solved, and high-precision and robust lung nodule segmentation are achieved.

CN120543576APending Publication Date: 2025-08-26HAINAN UNIV
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510617171.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The prior art faces the problems of target boundary accuracy, insufficient segmentation ability of small target areas, insufficient fusion of shallow and deep information, scarce medical image data and high labeling costs in the segmentation of lung nodules.

Method used

A potential spatial fusion method based on anatomical constraints is adopted, combined with the KL divergence constraint strategy, a multimodal collaborative optimization FSUNet architecture is built, and a cross-scale feature pyramid fusion mechanism and channel attention gating are used to design a Dice-Focal joint loss function, and the forward diffused noise sample generation strategy is used to suppress complex background noise interference.

Benefits of technology

It significantly improves the accuracy and robustness of lung nodule segmentation, can capture lung nodule boundaries and morphological characteristics more accurately, reduces missegment, enhances the recognition ability of small-volume and low-contrast nodules, and maintains stability in complex backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543576A_ABST
    Figure CN120543576A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical treatment, and discloses a pulmonary nodule segmentation method and system based on a diffusion model, and the method employs a potential space fusion method based on anatomical constraint, integrates the anatomical priori knowledge of lung tissue contour and the like through a variational auto-encoder (VAE), and combines a KL divergence constraint strategy. And the sensitivity of the model to a tiny focus is effectively improved. Secondly, a multi-modal collaborative optimization FSUNet architecture is constructed, and a cross-scale feature pyramid fusion mechanism is utilized to cooperate with a channel attention gating and dynamic noise scheduling algorithm, so that the expression ability of multi-dimensional features is enhanced. And finally, designing a Dice-Focal joint loss function to suppress complex background noise interference, and meanwhile, based on a forward diffusion-based noise sample generation strategy, remarkably reducing the dependence of the model on annotated data. Experimental results on an LIDC-IDRI data set show that the diffusion segmentation method based on multi-scale feature fusion not only shows excellent performance, but also shows excellent performance on an FPS index at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical technology, and in particular relates to a lung nodule segmentation method and system based on a diffusion model. Background Art

[0002] Accurate segmentation of lung nodules is crucial for assisting doctors in early diagnosis and developing treatment plans. However, due to challenges such as large size variation, low contrast with surrounding tissue, and blurred boundaries, existing segmentation models often struggle to cope with these complexities, making medical image segmentation a long-standing challenge.

[0003] Early-stage lung cancer usually manifests as isolated small lung nodules with irregular shapes, random distribution, and low contrast with surrounding normal tissue. These characteristics make doctors face greater subjectivity in imaging judgments, and also increase the difficulty and complexity of early diagnosis. Therefore, accurate lung nodule segmentation technology plays a vital role in early diagnosis. It not only provides auxiliary information to help doctors improve diagnostic accuracy, but also optimizes treatment plans and improves patient prognosis. However, due to the significant size variation of lung nodules in images, blurred boundaries and complex backgrounds, the lung nodule segmentation task has always been fraught with challenges, and traditional segmentation methods are difficult to meet actual clinical needs.

[0004] In recent years, deep learning technology has made significant progress in the field of medical image segmentation. Unlike traditional feature extraction methods that rely on manual experience, deep learning models significantly improve segmentation efficiency and accuracy through automatic feature extraction. Deep learning has particularly brought breakthroughs to medical image processing in complex tasks such as pulmonary nodule segmentation. Existing technologies have proposed a pulmonary nodule segmentation model based on a combination of a ResNet-50 encoder and multi-layer feature decoding technology, enhancing feature reuse and boundary capture capabilities, thereby significantly improving segmentation accuracy. Existing technologies have improved the U-Net model, combining multi-scale feature fusion with a lung lobe segmentation module, significantly improving pulmonary nodule segmentation accuracy and detection capabilities. Existing technologies have proposed the Wavelet U-Net++ model, which enhances multi-scale feature extraction through wavelet transforms, further improving the segmentation of small nodules. Existing technologies combine attention mechanisms with edge optimization algorithms, particularly focusing on cell edges, significantly improving segmentation accuracy, and achieving excellent results in cell morphology analysis tasks. These methods, through innovative network architectures and enhanced feature extraction capabilities, have promoted progress in early disease detection and diagnosis. However, when it comes to lung nodule segmentation, these methods still have shortcomings in capturing details, segmenting small target areas, and fusion of scale features, and urgently need further optimization.

[0005] In addition to technical challenges, medical image segmentation also faces challenges such as data scarcity and high annotation costs. Medical image annotation relies heavily on specialized knowledge, and subjective differences between experts and inconsistent annotation quality make deep learning segmentation methods less robust when working with small datasets. This is particularly true when processing complex lung nodule images, where the accuracy and stability of existing segmentation methods still have significant room for improvement.

[0006] Against this backdrop, diffusion models, as an emerging generative model, have demonstrated significant potential in the field of medical image generation and segmentation in recent years. Unlike traditional image segmentation methods that generate results in one go, diffusion models complete image generation through two stages: forward diffusion and backward diffusion. In the forward diffusion stage, Gaussian noise is gradually added to the original image until it approaches a Gaussian distribution. In the backward diffusion stage, a trained denoising deep network is used to gradually denoise the image, restoring a clear image. The gradual denoising nature of this process gives diffusion models significant advantages in capturing target details, particularly in complex backgrounds and capturing details, demonstrating greater robustness. Existing technologies have employed semantic 3D medical image synthesis methods based on conditional diffusion models, effectively improving the segmentation accuracy of 3D medical images. By generating high-quality segmented images through diffusion models, existing technologies successfully address the label scarcity issue and achieve accurate semantic segmentation across multiple medical image datasets. Existing technologies have proposed implicit image segmentation ensemble methods based on diffusion models. By integrating multiple diffusion models, these methods improve the accuracy and robustness of chest CT and brain MRI image segmentation. However, although the diffusion model has shown great potential in medical image segmentation, existing models still face some challenges in the task of lung nodule segmentation, especially in terms of target boundary accuracy, segmentation ability of small target areas, and fusion of shallow and deep information.

[0007] Through the above analysis, the problems and defects of the existing technology are as follows:

[0008] Although diffusion models have shown great potential in medical image segmentation, existing models still face some challenges in the task of lung nodule segmentation, especially in terms of target boundary accuracy, segmentation ability of small target areas, and fusion of shallow and deep information. Summary of the Invention

[0009] In response to the problems existing in the prior art, the present invention provides a lung nodule segmentation method and system based on a diffusion model.

[0010] The present invention is implemented as follows: a lung nodule segmentation method and system based on a diffusion model includes:

[0011] Step 1: A latent space fusion method based on anatomical constraints is used to integrate anatomical prior knowledge such as lung tissue contours through a variational autoencoder and combined with a KL divergence constraint strategy;

[0012] Step 2: Build a multimodal collaborative optimization FSUNet architecture, using a cross-scale feature pyramid fusion mechanism, combined with channel attention gating and dynamic noise scheduling algorithms;

[0013] Step 3: Design the Dice-Focal joint loss function to suppress complex background noise interference, and at the same time generate noise samples based on the forward diffusion strategy.

[0014] Furthermore, the forward diffusion:

[0015] The forward diffusion process simulates image degradation and gradually adds Gaussian noise to the initial latent variable zp (usually generated by the input image encoding) through a Markov chain, generating a series of noised states until zT; the process can be formally expressed as:

[0016]

[0017] in, is the cumulative noise scheduling parameter, α s =1-β s ,β s ∈(0,1) is the predefined noise intensity, t represents the time step, ∈ t is the random noise sampled from the standard normal distribution N(0,I); the noise intensity β s Usually a preset linear or cosine scheduling strategy is adopted;

[0018] Backward denoising is the inverse process of forward diffusion, and its goal is to gradually recover the original latent variables from the highly noisy state zT FSUNet, as a denoising network, receives the current diffusion state zt and time step information t as input and predicts noise based on the deep learning model. The denoising process updates the latent variables by the following formula:

[0019]

[0020] Among them, σ t is the posterior variance, usually set to σ t =β t , η is random noise, which is used to enhance sampling diversity; FSUNet extracts multi-scale features through the encoder, combines with SENet to dynamically adjust channel weights to focus on key information, and uses FPN to fuse shallow details with deep semantics to optimize noise prediction accuracy; the denoising process iterates from t = T to t = 0, and finally generates the denoised latent variable

[0021] Decoding and segmentation mask generation

[0022] Denoised latent variables The input pre-trained VAE decoder Ds is mapped to a feature representation with the same spatial dimensions as the input image through multi-layer deconvolution operations; the decoding process can be expressed as:

[0023]

[0024] in, Based on the recovered image latent representation, a lung nodule segmentation mask Mnodule is subsequently generated through thresholding or post-processing steps.

[0025] Furthermore, the complex background noise interference is suppressed:

[0026] Anatomy Controller Module

[0027] 1) Input processing and anatomical supervision

[0028] The Anatomy Controller module takes the original lung nodule image x as input and introduces the anatomical mask ma as a supervisory signal. The anatomical mask ma contains information such as lung contours, tracheal position, and key vascular structures, aiming to enhance the model's understanding and learning of lung anatomical features.

[0029] 2) Encoding stage: latent representation generation

[0030] In the encoding stage, the module uses the pre-trained VAE encoder Ev to extract features from the input image x; through multi-layer convolution operations and nonlinear activation functions, the encoder gradually compresses the high-dimensional image information into the parameters of the potential distribution, that is, the mean μ a and log variance The process can be expressed as:

[0031]

[0032] Subsequently, the latent variable za follows a multivariate normal distribution To ensure the differentiability of the training process, the reparameterization technique is used to calculate za:

[0033] z a =μ a +∈ a ⊙σ a (5)

[0034] Where ∈a is random noise sampled from a standard normal distribution, and ⊙ represents element-wise multiplication. This process not only preserves the semantic information of the image but also enhances the model’s ability to model potential variations in anatomical structures (such as morphological diversity or boundary ambiguity) through the introduction of random sampling.

[0035] 3) Decoding stage: latent space mapping and noise injection

[0036] z p =D v (z a )+F(z a )(6)

[0037] In the decoding stage, the latent variable za is input to the VAE decoder Dv; the decoder maps the high-dimensional latent representation za to a low-dimensional latent variable zp through multiple layers of deconvolution and upsampling operations, providing input for the subsequent FSUNet denoising network; in this process, the supervisory signal of the anatomical mask ma is embedded in the training target of the decoder to guide the model to learn the relevant features of the anatomical structure; the latent space mapping process can be formalized as:

[0038] Where F(·) is a nonlinear mapping function used to adjust the uniformity and expressiveness of the potential space distribution. To further enhance the robustness of the potential representation and adapt to complex segmentation scenarios, random Gaussian noise is introduced in the mapping process:

[0039]

[0040] Here, σ n is the noise intensity hyperparameter; noise injection not only enriches the diversity of the latent space, but also simulates the noise interference that may exist in real lung images, thereby improving the generalization ability of the model;

[0041] 4) Loss function optimization

[0042] Design a comprehensive loss function including reconstruction loss and KL divergence loss; reconstruction loss Lrecon is used to measure the decoder reconstructed image The similarity between the original image x is defined using the mean square error (MSE):

[0043]

[0044] KL divergence loss LKL constrains the latent variable distribution q(z a |x) is close to the prior distribution N(0,I) to improve the regularity and generalization ability of the latent space:

[0045]

[0046] Where d is the dimension of the latent space; the comprehensive loss function is:

[0047] A_Loss=L recon +λ·L KL (10)

[0048] The weight parameter λ balances the relationship between reconstruction accuracy and regularity of the latent space. Through multiple rounds of experimental evaluation, it was found that λ = 0.5 can achieve better performance between reconstruction quality and model stability. To further verify its impact, subsequent ablation experiments will analyze the specific contribution of different λ values ​​to segmentation performance. During training, the Adam optimizer is used (with an initial learning rate of 10-3).

[0049] Furthermore, the FSUNet architecture:

[0050] 1) VAE-guided information fusion

[0051] The pre-trained Es provides prior knowledge of the lung anatomical structure by learning the potential distribution of the image, effectively guiding the denoising and restoration process. Specifically, the latent variable zs output by Es is fused with the intermediate feature map of the encoder through a skip connection, which can be mathematically expressed as:

[0052]

[0053] Among them, F enc is the feature map extracted by the encoder, ⊕ represents the feature fusion operation, z s ∈R d is the latent variable, and d is the latent space dimension. This process not only enhances the model’s sensitivity to lung nodule details, but also improves the stability of the denoising process.

[0054] 2) Adaptive SE module

[0055] FSUNet enhances its ability to capture key features of lung nodules by embedding the SENet attention mechanism. In the two-stage encoder-decoder architecture, the network cascades SENet units after each 3×3 convolutional module, achieving adaptive recalibration of feature channels through the channel attention mechanism.

[0056] 3) Multi-scale FPN module;

[0057] 4) Loss function.

[0058] Furthermore, the multi-scale FPN module

[0059] The top-down pathway aims to propagate high-level semantic information from deep feature maps back to shallow scales and gradually restore spatial resolution through upsampling operations. Starting from E5, the nearest neighbor interpolation method is used to upsample deep feature maps to align the spatial dimensions with those of shallow feature maps. This process ensures the effective integration of deep semantic information and shallow spatial details, providing support for subsequent fusion operations. The upsampling operation can be expressed as:

[0060] Ui=Upsample(E(i+1),scale=2) (12)

[0061] Where Ui represents the upsampled feature map, Upsample is the upsampling function, and scale = 2 means the resolution is magnified by a factor of 2. The top-down path starts from E5 and generates upsampled feature maps U4 to U1 through the green dotted arrows marked "2xUpsample". These feature maps are aligned with the corresponding Ei feature maps in the spatial dimension, creating conditions for the fusion operation of the horizontal connection.

[0062] Lateral connections generate multi-scale feature maps by fusing the features of the bottom-up and top-down pathways, further improving the model's ability to perceive lung nodules of different sizes. Specifically, for each layer Ei (i = 1, 2, ..., 5), a 1×1 convolution is first used to adjust its channel number to be consistent with the corresponding upsampled feature Ui. Subsequently, fusion is completed through element-wise addition. The fusion process can be expressed as:

[0063] Fi'=Ui+Conv 1×1 (Ei)

[0064] Fi'=Concat(Ui,Conv 1×1 (Ei)) (13)

[0065] Where Fi' is the fused feature map, Conv1×1 represents a 1×1 convolution operation for channel alignment, and Concat represents feature concatenation along the channel axis. The fused Fi' contains both deep semantic information and shallow spatial details, significantly improving the expressiveness of features. These feature maps are ultimately fed into the FSUNet decoder to generate accurate segmentation predictions.

[0066] Furthermore, the loss function:

[0067] The loss function is crucial in deep learning models by quantifying the difference between the predicted value and the true label; its mathematical definition is:

[0068]

[0069] Among them, Pi represents the pixel probability predicted by the model, ti is the pixel value of the true label, N is the total number of pixels, ε DL is a smoothing factor to avoid zero denominator;

[0070] This paper introduces Focal Loss to enhance the model's ability to focus on small lesions and low-contrast areas by dynamically adjusting the weights of difficult-to-classify samples. The formula is:

[0071] FocalLoss=-α FL (1-p t ) γ log(p t ) (15)

[0072] Among them, pt is the predicted probability of the target category, α FL is a balancing factor (set to 0.25 in this paper), and γ is an adjustment factor (set to 2) used to control the weighting degree of difficult-to-classify samples; FocalLoss reduces the loss contribution of easy-to-classify samples, allowing the model to focus on small nodules and boundary areas that are difficult to segment;

[0073] The designed joint loss function integrates the advantages of DiceLoss and FocalLoss, and its form is:

[0074] L loss =DiceLoss+τ·FocalLoss (16)

[0075] Among them, τ=1 is the weight factor, which has been experimentally verified to effectively balance the contributions of the two losses; DiceLoss ensures the accuracy of the overall segmentation area, while FocalLoss improves the attention to difficult samples. The synergistic effect of the two significantly enhances the performance of the FSMedDiff model.

[0076] Another object of the present invention is to provide a pulmonary nodule segmentation system based on a diffusion model, comprising:

[0077] The fusion module is used to adopt an anatomically constrained latent space fusion method, integrating anatomical prior knowledge such as lung tissue contours through a variational autoencoder and combining it with a KL divergence constraint strategy;

[0078] A building block for constructing the FSUNet architecture for multimodal collaborative optimization, utilizing a cross-scale feature pyramid fusion mechanism, coupled with channel attention gating and dynamic noise scheduling algorithms;

[0079] The suppression module is used to design the Dice-Focal joint loss function to suppress complex background noise interference, while also generating noise samples based on the forward diffusion strategy.

[0080] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the lung nodule segmentation method based on the diffusion model.

[0081] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the lung nodule segmentation method based on the diffusion model.

[0082] Another object of the present invention is to provide an information data processing terminal, which is used to implement the lung nodule segmentation system based on the diffusion model.

[0083] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are analyzed from the following aspects:

[0084] First, in view of the technical problems existing in the above-mentioned prior art and the difficulty of solving these problems, this paper closely combines the technical solutions to be protected by the present invention and the results and data during the research and development process, and analyzes in detail and in depth how the technical solutions of the present invention solve the technical problems and some creative technical effects brought about by solving the problems. The specific description is as follows:

[0085] In the field of medical image segmentation, lung nodule segmentation faces many challenges that are difficult to effectively address with existing technologies. This paper proposes innovative solutions from multiple key aspects, significantly improving the performance of lung nodule segmentation and achieving outstanding technical results.

[0086] Existing technical problems and difficulty in solving them: Pulmonary nodules vary greatly in size, have low contrast with surrounding tissues, and have blurred boundaries. Medical image annotation data is scarce, the annotation cost is high, and there are subjective differences in annotations by different experts. These factors make the task of pulmonary nodule segmentation extremely challenging. Traditional segmentation methods rely on manual experience to extract features and perform poorly in complex pulmonary nodule segmentation tasks. Although deep learning methods have made progress, they still have shortcomings in detail capture, small target area segmentation, and scale feature fusion. Although the diffusion model has potential, it also needs to be improved in terms of target boundary accuracy, small target area segmentation capabilities, and shallow and deep information fusion in pulmonary nodule segmentation.

[0087] Technical solution and solution of the present invention:

[0088] Latent space fusion based on anatomical constraints: A variational autoencoder is used to integrate anatomical prior knowledge, such as lung tissue contours, and combined with a KL divergence constraint strategy. Through the Anatomy Controller module, an anatomical mask is introduced as a supervisory signal to extract the parameters of the image's latent distribution during the encoding phase. Reparameterization techniques are used to enhance the model's ability to model anatomical variations. During the decoding phase, latent variables are mapped and noise is injected to enrich the latent space diversity. A comprehensive loss function, consisting of a reconstruction loss and a KL divergence loss, is designed to optimize the embedding and utilization of anatomical information and improve the model's sensitivity to subtle lesions.

[0089] The FSUNet architecture for multimodal collaborative optimization is constructed: Utilizing a cross-scale feature pyramid fusion mechanism, coupled with channel attention gating and a dynamic noise scheduling algorithm, the VAE guides information fusion, fusing the latent variables of the pre-trained VAE encoder with the encoder's intermediate feature maps to enhance sensitivity to lung nodule details and denoising stability. The adaptive SE module uses a channel attention mechanism to adaptively recalibrate feature channels, strengthening the ability to capture key features. The multi-scale FPN module transmits deep semantic information through a top-down pathway, horizontally connecting and fusing multi-scale features to improve the perception of lung nodules of different sizes.

[0090] A joint Dice-Focal loss function is designed to suppress complex background noise interference while also generating noise samples based on forward diffusion. Dice Loss ensures the accuracy of the overall segmentation region, while Focal Loss dynamically adjusts the weights of difficult-to-classify samples to enhance the focus on small lesions and low-contrast areas. The two work together to improve model performance.

[0091] Creative technical effects:

[0092] Significantly Improved Segmentation Accuracy: On the LIDC-IDRI dataset, the proposed method achieved a Dice Similarity Coefficient (DSC) of 0.915, significantly outperforming traditional models such as ResNet-50 and U-Net++. The proposed method achieved a precision of 0.973, an average precision (AP) of 0.907, and a mean Intersection over Union (mIoU) of 0.953, demonstrating its ability to more accurately capture lung nodule boundaries and morphological features, reduce missegmentation, and enhance recognition of small, low-contrast nodules at varying confidence thresholds, achieving competitive pixel-level segmentation accuracy.

[0093] Enhanced robustness and generalization: By incorporating anatomical prior knowledge, optimizing multi-scale feature fusion, and designing a joint loss function, the model maintains stable performance in complex backgrounds and under diverse data conditions. Ablation experiments demonstrate that the collaborative work of the AnatomyController module, the SENet attention module, and the multi-scale FPN module significantly improves model performance and enhances its adaptability to lung nodules of varying shapes, sizes, and locations.

[0094] Balancing precision and efficiency: With an inference speed of 35.5 FPS, it is at the Pareto frontier in the precision-speed trade-off diagram. This meets real-time requirements without significantly sacrificing segmentation performance, making it suitable for actual clinical application scenarios. It can assist doctors in quickly locating nodules, improving the efficiency and quality of clinical diagnosis.

[0095] The diffusion model can effectively capture the detailed information in medical images through the generation process of gradual denoising, thereby improving the segmentation performance. The present invention constructs the FSMedDiff diffusion model. The model achieves accurate segmentation by fusing the Anatomy Controller module and the FSUNet architecture. The Anatomy Controller module embeds the lung anatomy prior information into the latent space to enhance the lesion perception ability; the FSUNet architecture integrates the multi-scale feature pyramid and the channel attention mechanism, and adopts a dynamic noise scheduling strategy to achieve collaborative optimization of cross-scale features; at the same time, a composite loss function is designed to optimize the global segmentation accuracy while strengthening the focusing ability on the small nodule area, significantly improving the segmentation robustness in complex noise environments. Experiments on the LIDC-IDRI dataset have shown that the segmentation Dice similarity coefficient (DSC) of FSMedDiff reaches 0.915, and the inference speed is 35.5FPS, which not only shows excellent segmentation accuracy, but also has high real-time performance.

[0096] This paper proposes a diffusion segmentation model (FSMedDiff) based on multi-scale feature fusion. The model optimizes multi-scale feature fusion through feature pyramid network (FPN)[8] and channel attention mechanism (SENet)[9], introduces the Anatomy Controller module to embed anatomical prior information, and uses the generation capability of the diffusion model to alleviate the data scarcity problem, thereby significantly improving the accuracy, robustness and generalization ability of pulmonary nodule segmentation.

[0097] This paper first analyzes the challenges faced by the lung nodule segmentation task in dealing with multi-scale targets, fuzzy boundaries, complex backgrounds, and data annotation scarcity. An innovative medical image segmentation framework FSMedDiff is proposed, which is mainly improved in three key technical aspects. First, a latent space fusion method based on anatomical constraints is adopted to integrate anatomical prior knowledge such as lung tissue contours through a variational autoencoder (VAE), and combined with the KL divergence constraint strategy, the sensitivity of the model to small lesions is effectively improved. Secondly, a multimodal collaborative optimization FSUNet architecture is constructed, and the cross-scale feature pyramid fusion mechanism is used, combined with channel attention gating and dynamic noise scheduling algorithm to enhance the expression ability of multi-dimensional features. Finally, the Dice-Focal joint loss function is designed to suppress the interference of complex background noise. At the same time, the noise sample generation strategy based on forward diffusion significantly reduces the model's dependence on labeled data. Experimental results on the LIDC-IDRI dataset show that the diffusion segmentation method with multi-scale feature fusion not only exhibits excellent performance, but also performs well in the FPS indicator. To further verify the effectiveness and rationality of the method, the author also conducted detailed ablation experiments and qualitative analysis, which fully demonstrated that the method effectively improved model performance and demonstrated excellent generalization ability without increasing additional inference cost.

[0098] Second, as auxiliary evidence for the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:

[0099] The expected benefits and commercial value of the technical solution of the present invention after conversion are:

[0100] Expected Returns: Considering the current market growth in assisted diagnosis, telemedicine, and health management, this technology is expected to generate product and service sales revenue exceeding 200 million RMB within three years, with a compound annual growth rate exceeding 30%. Leveraging its technological advantages and differentiated capabilities, the project's overall return on investment is projected to exceed 50%, achieving break-even by the end of the second year. The improved diagnostic and R&D efficiency achieved by this system for medical institutions and pharmaceutical companies will save partners over 10 million RMB in operating and R&D costs annually.

[0101] The lung nodule segmentation technology based on the diffusion model has the following three core values:

[0102] 1. Expansion of the auxiliary diagnosis market

[0103] With its highly accurate nodule boundary annotation, this system significantly shortens image reading time, reduces misdiagnoses and missed diagnoses, and helps hospitals optimize processes and enhance the patient experience. With rising lung cancer incidence and increased health awareness, this technology is expected to attract procurement from medical institutions, rapidly capture the auxiliary diagnosis market, and generate stable sales revenue.

[0104] 2. Improve drug development efficiency

[0105] Accurate quantitative analysis of nodule volume and morphology can provide reliable data for dose design and efficacy evaluation of new drug candidates, accelerating the screening and optimization process, reducing R&D costs, and shortening time to market. In targeted lung cancer drug research, real-time monitoring of the drug's impact on nodule changes can help identify the mechanism and improve the success rate of R&D.

[0106] 3. Support telemedicine and health management

[0107] This technology can be seamlessly integrated into remote diagnostic platforms, enabling online automated segmentation and initial screening reports, expanding healthcare coverage and reducing operating costs. It also provides regular lung monitoring for high-risk individuals, capturing lesion dynamics in a timely manner and facilitating personalized risk assessment and intervention. Medical institutions and health management companies can leverage this to expand value-added services, enhance user engagement, and open up new revenue streams. BRIEF DESCRIPTION OF THE DRAWINGS

[0108] Figure 1 This is a flow chart of a lung nodule segmentation method based on a diffusion model provided in an embodiment of the present invention.

[0109] Figure 2 This is a structural block diagram of a lung nodule segmentation system based on a diffusion model provided in an embodiment of the present invention.

[0110] Figure 3 This is a structural diagram of the FSMedDiff overall model provided by an embodiment of the present invention.

[0111] Figure 4 This is a module structure diagram of the Anatomy Controller provided in an embodiment of the present invention.

[0112] Figure 5 This is a diagram of the FSUNet network structure provided by an embodiment of the present invention.

[0113] Figure 6 It is a structural diagram of the adaptive SE module provided by an embodiment of the present invention.

[0114] Figure 7 4 is a structural diagram of a multi-scale FPN provided by an embodiment of the present invention.

[0115] Figure 8 This is a trend diagram of the training loss function provided by an embodiment of the present invention.

[0116] Figure 9 This is a model accuracy-speed trade-off diagram provided by an embodiment of the present invention.

[0117] Figure 10 This is a segmentation visualization result diagram provided by an embodiment of the present invention.

[0118] Figure 11 This is a comparison chart of Precision, AP, and mIoU provided by an embodiment of the present invention.

[0119] Figure 12 It is a DSC comparison chart provided by an embodiment of the present invention.

[0120] Figure 13 This is a comparison chart of inference speed provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0121] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0122] like Figure 1 As shown, the embodiment of the present invention provides a lung nodule segmentation method and system based on a diffusion model, which includes the following steps:

[0123] S101 uses a latent space fusion method based on anatomical constraints, integrates anatomical prior knowledge such as lung tissue contours through a variational autoencoder, and combines it with a KL divergence constraint strategy;

[0124] S102: Build a multimodal collaborative optimization FSUNet architecture, using a cross-scale feature pyramid fusion mechanism, combined with channel attention gating and dynamic noise scheduling algorithms;

[0125] S103, design the Dice-Focal joint loss function to suppress complex background noise interference, and at the same time generate noise samples based on the forward diffusion strategy.

[0126] The forward diffusion provided by the embodiment of the present invention:

[0127] The forward diffusion process simulates image degradation and gradually moves to the initial latent variable z through the Markov chain. p Add Gaussian noise (usually generated by encoding the input image) to generate a series of noise states up to z T ; The process can be formally expressed as:

[0128]

[0129] in, is the cumulative noise scheduling parameter, α s =1-β s ,β s ∈(0,1) is the predefined noise intensity, t represents the time step, ∈ t is the random noise sampled from the standard normal distribution N(0,I); the noise intensity β sUsually a preset linear or cosine scheduling strategy is adopted;

[0130] Backward denoising is the inverse process of forward diffusion, and its goal is to gradually recover the original latent variables from the highly noisy state zT FSUNet, as a denoising network, receives the current diffusion state zt and time step information t as input and predicts noise based on the deep learning model. The denoising process updates the latent variables by the following formula:

[0131]

[0132] Among them, σ t is the posterior variance, usually set to σ t =β t , η is random noise, which is used to enhance sampling diversity; FSUNet extracts multi-scale features through the encoder, combines with SENet to dynamically adjust channel weights to focus on key information, and uses FPN to fuse shallow details with deep semantics to optimize noise prediction accuracy; the denoising process iterates from t = T to t = 0, and finally generates the denoised latent variable

[0133] Decoding and segmentation mask generation

[0134] Denoised latent variables The input pre-trained VAE decoder Ds is mapped to a feature representation with the same spatial dimensions as the input image through multi-layer deconvolution operations; the decoding process can be expressed as:

[0135]

[0136] in, Based on the recovered image latent representation, a lung nodule segmentation mask Mnodule is subsequently generated through thresholding or post-processing steps.

[0137] The embodiment of the present invention provides the following method for suppressing complex background noise interference:

[0138] Anatomy Controller Module

[0139] 1) Input processing and anatomical supervision

[0140] The Anatomy Controller module takes the original lung nodule image x as input and introduces the anatomical mask ma as a supervisory signal. The anatomical mask ma contains information such as lung contours, tracheal position, and key vascular structures, aiming to enhance the model's understanding and learning of lung anatomical features.

[0141] 2) Encoding stage: latent representation generation

[0142] In the encoding stage, the module uses the pre-trained VAE encoder Ev to extract features from the input image x; through multi-layer convolution operations and nonlinear activation functions, the encoder gradually compresses the high-dimensional image information into the parameters of the potential distribution, that is, the mean μ a and log variance The process can be expressed as:

[0143]

[0144] Subsequently, the latent variable za follows a multivariate normal distribution To ensure the differentiability of the training process, the reparameterization technique is used to calculate za:

[0145] z a =μ a +∈ a ⊙σ a (5)

[0146] Where ∈a is random noise sampled from a standard normal distribution, and ⊙ represents element-wise multiplication. This process not only preserves the semantic information of the image but also enhances the model’s ability to model potential variations in anatomical structures (such as morphological diversity or boundary ambiguity) through the introduction of random sampling.

[0147] 3) Decoding stage: latent space mapping and noise injection

[0148] z p =D v (z a )+F(z a )(6)

[0149] In the decoding stage, the latent variable za is input to the VAE decoder Dv; the decoder maps the high-dimensional latent representation za to a low-dimensional latent variable zp through multiple layers of deconvolution and upsampling operations, providing input for the subsequent FSUNet denoising network; in this process, the supervisory signal of the anatomical mask ma is embedded in the training target of the decoder to guide the model to learn the relevant features of the anatomical structure; the latent space mapping process can be formalized as:

[0150] Where F(·) is a nonlinear mapping function used to adjust the uniformity and expressiveness of the potential space distribution. To further enhance the robustness of the potential representation and adapt to complex segmentation scenarios, random Gaussian noise is introduced in the mapping process:

[0151]

[0152] Here, σ n is the noise intensity hyperparameter; noise injection not only enriches the diversity of the latent space, but also simulates the noise interference that may exist in real lung images, thereby improving the generalization ability of the model;

[0153] 4) Loss function optimization

[0154] Design a comprehensive loss function including reconstruction loss and KL divergence loss; reconstruction loss Lrecon is used to measure the decoder reconstructed image The similarity between the original image x is defined using the mean square error (MSE):

[0155]

[0156] KL divergence loss LKL constrains the latent variable distribution q(z a |x) is close to the prior distribution N(0,I) to improve the regularity and generalization ability of the latent space:

[0157]

[0158] Where d is the dimension of the latent space; the comprehensive loss function is:

[0159] A_Loss=L recon +λ·L KL (10)

[0160] The weight parameter λ balances the relationship between reconstruction accuracy and regularity of the latent space. Through multiple rounds of experimental evaluation, it was found that λ = 0.5 can achieve better performance between reconstruction quality and model stability. To further verify its impact, subsequent ablation experiments will analyze the specific contribution of different λ values ​​to segmentation performance. During training, the Adam optimizer is used (with an initial learning rate of 10-3).

[0161] The FSUNet architecture provided by the embodiment of the present invention:

[0162] 1) VAE-guided information fusion

[0163] The pre-trained Es provides prior knowledge of the lung anatomical structure by learning the potential distribution of the image, effectively guiding the denoising and restoration process. Specifically, the latent variable zs output by Es is fused with the intermediate feature map of the encoder through a skip connection, which can be mathematically expressed as:

[0164]

[0165] Among them, F enc is the feature map extracted by the encoder, ⊕ represents the feature fusion operation, z s ∈R d is the latent variable, and d is the latent space dimension. This process not only enhances the model’s sensitivity to lung nodule details, but also improves the stability of the denoising process.

[0166] 2) Adaptive SE module

[0167] FSUNet enhances its ability to capture key features of lung nodules by embedding the SENet attention mechanism. In the two-stage encoder-decoder architecture, the network cascades SENet units after each 3×3 convolutional module, achieving adaptive recalibration of feature channels through the channel attention mechanism.

[0168] 3) Multi-scale FPN module

[0169] The top-down pathway aims to propagate high-level semantic information from deep feature maps back to shallow scales and gradually restore spatial resolution through upsampling operations. Starting from E5, the nearest neighbor interpolation method is used to upsample deep feature maps to align the spatial dimensions with those of shallow feature maps. This process ensures the effective integration of deep semantic information and shallow spatial details, providing support for subsequent fusion operations. The upsampling operation can be expressed as:

[0170] Ui=Upsample(E(i+1),scale=2) (12)

[0171] Where Ui represents the upsampled feature map, Upsample is the upsampling function, and scale = 2 means the resolution is magnified by 2 times; Figure 7 As shown in the figure, the top-down path starts from E5 and generates upsampled feature maps U4 to U1 through the green dotted arrows marked as "2xUpsample". These feature maps are aligned with the corresponding Ei feature maps in the spatial dimension, creating conditions for the fusion operation of the horizontal connection.

[0172] Lateral connections generate multi-scale feature maps by fusing the features of the bottom-up and top-down pathways, further improving the model's ability to perceive lung nodules of different sizes. Specifically, for each layer Ei (i = 1, 2, ..., 5), a 1×1 convolution is first used to adjust its channel number to be consistent with the corresponding upsampled feature Ui. Subsequently, fusion is completed through element-wise addition. The fusion process can be expressed as:

[0173] Fi'=Ui+Conv 1×1 (Ei)

[0174] Fi'=Concat(Ui,Conv 1×1 (Ei)) (13)

[0175] Where Fi' is the fused feature map, Conv1×1 represents a 1×1 convolution operation for channel alignment, and Concat represents feature concatenation along the channel axis. The fused Fi' contains both deep semantic information and shallow spatial details, significantly improving the expressiveness of features. These feature maps are ultimately fed into the FSUNet decoder to generate accurate segmentation predictions.

[0176] 4) Loss Function

[0177] The loss function is crucial in deep learning models by quantifying the difference between the predicted value and the true label; its mathematical definition is:

[0178]

[0179] Among them, Pi represents the pixel probability predicted by the model, ti is the pixel value of the true label, N is the total number of pixels, ε DL is a smoothing factor to avoid zero denominator;

[0180] This paper introduces FocalLoss to enhance the model's ability to focus on small lesions and low-contrast areas by dynamically adjusting the weights of difficult-to-classify samples. The formula is:

[0181] FocalLoss=-α FL (1-p t ) γ log(p t ) (15)

[0182] Among them, pt is the predicted probability of the target category, α FL is a balancing factor (set to 0.25 in this paper), and γ is an adjustment factor (set to 2) used to control the weighting degree of difficult-to-classify samples; FocalLoss reduces the loss contribution of easy-to-classify samples, allowing the model to focus on small nodules and boundary areas that are difficult to segment;

[0183] The designed joint loss function integrates the advantages of DiceLoss and FocalLoss, and its form is:

[0184] L loss =DiceLoss+τ·FocalLoss (16)

[0185] Among them, τ=1 is the weight factor, which has been experimentally verified to effectively balance the contributions of the two losses; DiceLoss ensures the accuracy of the overall segmentation area, while FocalLoss improves the attention to difficult samples. The synergistic effect of the two significantly enhances the performance of the FSMedDiff model.

[0186] like Figure 2As shown, an embodiment of the present invention provides a lung nodule segmentation system based on a diffusion model, including:

[0187] The fusion module is used to adopt an anatomically constrained latent space fusion method, integrating anatomical prior knowledge such as lung tissue contours through a variational autoencoder and combining it with a KL divergence constraint strategy;

[0188] A building block for constructing the FSUNet architecture for multimodal collaborative optimization, utilizing a cross-scale feature pyramid fusion mechanism, coupled with channel attention gating and dynamic noise scheduling algorithms;

[0189] The suppression module is used to design the Dice-Focal joint loss function to suppress complex background noise interference, while also generating noise samples based on the forward diffusion strategy.

[0190] Another object of the present invention is to provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the lung nodule segmentation method based on the diffusion model.

[0191] Another object of the present invention is to provide a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the lung nodule segmentation method based on the diffusion model.

[0192] Another object of the present invention is to provide an information data processing terminal, which is used to implement the lung nodule segmentation system based on the diffusion model.

[0193] The present invention is specifically implemented:

[0194] 1. This paper proposes a diffusion segmentation model (FSMedDiff) based on multi-scale feature fusion. This model optimizes multi-scale feature fusion through the Feature Pyramid Network (FPN) and the Channel Attention Mechanism (SENet), introduces the AnatomyController module to embed anatomical prior information, and leverages the generative power of the diffusion model to alleviate the data scarcity problem, thereby significantly improving the accuracy, robustness, and generalization of pulmonary nodule segmentation.

[0195] 2 Network Structure

[0196] 2.1 Overall Architecture Design

[0197] The task of pulmonary nodule segmentation requires the model to accurately extract nodule regions from lung CT images to support clinical applications. However, the diversity of nodule sizes, fuzzy boundaries, and complex background interference pose challenges to traditional methods. To this end, FSMedDiff overcomes these difficulties through innovative design. Its overall architecture is as follows: Figure 3 As shown in the figure, it consists of the Anatomy Controller module and the FSUNet encoder-decoder network. The Anatomy Controller uses a variational autoencoder (VAE) to learn a latent representation of the input image, embedding anatomical prior information such as lung contours and tracheal position. This provides structured guidance in the latent space, ensuring anatomical consistency of the segmentation results and enhancing perception of low-contrast areas. This design is particularly critical for small object detection and background suppression.

[0198] Within the diffusion model framework, FSMedDiff employs a two-stage process consisting of forward diffusion and backward denoising. Forward diffusion simulates image degradation by gradually adding Gaussian noise, generating a highly noisy latent variable zT. Backward denoising removes the noise using FSUNet, restoring the original latent variable and generating a segmentation mask. Based on an encoder-decoder architecture, FSUNet integrates a pre-trained VAE encoder, SENet, and FPN to achieve multi-scale feature fusion and efficient denoising. Its design not only retains the generative advantages of the diffusion model but also improves its ability to capture small nodule details and robustness against complex backgrounds.

[0199] 2.1.2 Forward Diffusion Process

[0200] The forward diffusion process simulates image degradation and gradually adds Gaussian noise to the initial latent variable zp (usually generated by the input image encoding) through a Markov chain, generating a series of noised states until zT. The process can be formally expressed as:

[0201]

[0202] in, is the cumulative noise scheduling parameter, α s =1-β s ,β s ∈(0,1) is the predefined noise intensity, t represents the time step, ∈ t is the random noise sampled from the standard normal distribution N(0,I). Noise intensity β s A preset linear or cosine scheduling strategy is usually adopted to ensure that the noise increases gradually and simulate the continuity of the image degradation process. Through T steps of iteration, the initial clear latent variable zp is transformed into a highly noisy state zT, providing degraded samples for the subsequent denoising stage.

[0203] 2.1.3 Reverse denoising process

[0204] Backward denoising is the inverse process of forward diffusion, and its goal is to gradually recover the original latent variables from the highly noisy state zT FSUNet, as a denoising network, receives the current diffusion state zt and time step information t as input and predicts noise based on the deep learning model. The denoising process updates the latent variables by the following formula:

[0205]

[0206] Among them, σ t is the posterior variance, usually set to σ t =β t , η is random noise, used to enhance sampling diversity. FSUNet extracts multi-scale features through the encoder, combines with SENet to dynamically adjust channel weights to focus on key information, and uses FPN to fuse shallow details with deep semantics to optimize noise prediction accuracy. The denoising process iterates from t = T to t = 0, and finally generates the denoised latent variable

[0207] 2.1.4 Decoding and Segmentation Mask Generation

[0208] Denoised latent variables The input is a pre-trained VAE decoder Ds, which is mapped to a feature representation with the same spatial dimensions as the input image through multi-layer deconvolution operations. The decoding process can be expressed as:

[0209]

[0210] in, The recovered image latent representation is then used to generate a lung nodule segmentation mask Mnodule through thresholding or post-processing. This process combines anatomical priors with multi-scale feature fusion to ensure that the segmentation results are excellent in both anatomical consistency and boundary accuracy.

[0211] 2.2Anatomy Controller Module

[0212] In the task of pulmonary nodule segmentation, accurately identifying and segmenting small-sized nodule areas is a very challenging task, especially under complex lung background and fuzzy boundary conditions. Traditional models often produce mis-segmentation due to background noise interference and loss of high-frequency details. To address this problem, this paper proposes a novel Anatomy Controller module, which aims to provide structured guidance for subsequent segmentation tasks by learning anatomical prior information, thereby improving the model's perception of low-contrast areas and fuzzy boundaries. This module is based on the variational autoencoder (VAE) architecture design. By constructing a latent space representation of the input image, it captures the global characteristics of the lung anatomical structure. Its overall structure is as follows: Figure 4The module workflow consists of three main stages: input processing, latent representation generation, and decoding space mapping. It optimizes the embedding and utilization of anatomical information through a carefully designed loss function, ultimately providing high-quality latent representations for the FSUNet denoising network.

[0213] 2.2.1 Input Processing and Anatomy Supervision

[0214] The Anatomy Controller module takes the original image x of the lung nodule as input, and introduces the anatomical mask ma as a supervisory signal. The anatomical mask ma contains information such as the lung contour, tracheal position, and key vascular structures, aiming to enhance the model's understanding and learning ability of the anatomical features of the lung. Unlike the traditional method of directly using ma as the encoder input, this design indirectly guides the model to learn anatomical prior information in the latent space through a supervision mechanism during the training process. This strategy effectively avoids the information redundancy or overfitting problems that may be caused by directly splicing the mask, while ensuring that the latent space can efficiently represent the key characteristics of the anatomical structure. In the specific implementation, the input image x is first preprocessed (such as normalization operation) to meet the feature extraction requirements of the subsequent encoder.

[0215] 2.2.2 Encoding stage: latent representation generation

[0216] In the encoding stage, the module uses the pre-trained VAE encoder Ev to extract features from the input image x. Through multi-layer convolution operations and nonlinear activation functions, the encoder gradually compresses the high-dimensional image information into the parameters of the potential distribution, namely the mean μ a and log variance The process can be expressed as:

[0217]

[0218] Subsequently, the latent variable za follows a multivariate normal distribution To ensure the differentiability of the training process, the reparameterization technique is used to calculate za:

[0219] z a =μ a +∈ a ⊙σ a (5)

[0220] Where ∈a is random noise sampled from a standard normal distribution, and ⊙ represents element-wise multiplication. This process not only preserves the semantic information of the image, but also enhances the model’s ability to model potential variations in anatomical structures (such as morphological diversity or boundary ambiguity) through the introduction of random sampling.

[0221] 2.2.3 Decoding Stage: Latent Space Mapping and Noise Injection

[0222] zp =D v (z a )+F(z a )(6)

[0223] During the decoding phase, the latent variable za is input to the VAE decoder Dv. Through multiple layers of deconvolution and upsampling, the decoder maps the high-dimensional latent representation za into a low-dimensional latent variable zp, which provides input to the subsequent FSUNet denoising network. During this process, the supervisory signal of the anatomical mask ma is embedded in the decoder's training objective to guide the model in learning relevant features of the anatomical structure. The latent space mapping process can be formalized as:

[0224] Where F(·) is a nonlinear mapping function used to adjust the uniformity and expressiveness of the latent space distribution. To further enhance the robustness of the latent representation and adapt to complex segmentation scenarios, random Gaussian noise is introduced in the mapping process:

[0225]

[0226] Here, σ n is the noise intensity hyperparameter. Noise injection not only enriches the diversity of the latent space but also simulates the noise interference that may exist in real lung images, thereby improving the generalization ability of the model.

[0227] 2.2.4 Loss Function Optimization

[0228] To ensure the effectiveness of the Anatomy Controller module, a comprehensive loss function including reconstruction loss and KL divergence loss is designed. The reconstruction loss Lrecon is used to measure the decoder reconstructed image The similarity between the original image x is defined using the mean square error (MSE):

[0229]

[0230] KL divergence loss LKL constrains the latent variable distribution q(z a |x) is close to the prior distribution N(0,I) to improve the regularity and generalization ability of the latent space:

[0231]

[0232] Where d is the dimension of the latent space. The comprehensive loss function is:

[0233] A_Loss=L recon +λ·L KL (10)

[0234] The weight parameter λ balances reconstruction accuracy with latent space regularity. Through multiple rounds of experimental evaluation, we found that λ = 0.5 achieves an optimal balance between reconstruction quality and model stability. To further validate this impact, subsequent ablation experiments will analyze the specific contributions of different λ values ​​to segmentation performance. During training, the Adam optimizer (with an initial learning rate of 10-3) is used to iteratively update module parameters via mini-batch gradient descent to achieve rapid convergence and maximize optimization results.

[0235] 2.3FSUNet Codec Network

[0236] FSUNet is one of the core components of the FSMedDiff model, which aims to achieve accurate segmentation of lung nodules through image denoising and latent variable recovery. Figure 5 As shown in Figure 2, FSUNet adopts an encoder-decoder architecture and integrates advanced modules such as pre-trained variational autoencoders (ES), SENet, and feature pyramid networks. The synergy of these modules significantly improves the model's performance in multi-scale feature extraction, key information focusing, and complex background adaptation, thereby enhancing the accuracy and robustness of lung nodule segmentation.

[0237] 2.3.1 VAE-guided Information Fusion

[0238] To enhance the network's ability to learn the latent space, FSUNet incorporates the output of a pre-trained VAE encoder, Es, as conditional information. Unlike traditional feature concatenation, FSUNet incorporates the output of Es between the encoder and decoder via skip connections, providing rich context for the decoding process. This design enables the model to more precisely focus on lung nodule regions while reducing interference from background noise and redundant information.

[0239] The pre-trained Es provides prior knowledge of the lung anatomical structure by learning the potential distribution of the image, effectively guiding the denoising and restoration process. Specifically, the latent variable zs output by Es is fused with the intermediate feature map of the encoder through a skip connection, which can be mathematically expressed as:

[0240]

[0241] Among them, F enc is the feature map extracted by the encoder, ⊕ represents the feature fusion operation, z s ∈R d is the latent variable, and d is the latent space dimension. This process not only enhances the model’s sensitivity to lung nodule details but also improves the stability of the denoising process.

[0242] 2.3.2 Adaptive SE Module

[0243] FSUNet strengthens the ability to capture the key features of lung nodules by embedding SENet attention mechanism. The network is in a two-stage architecture of encoder and decoder, such as Figure 6 As shown, SENet units are cascaded after each 3×3 convolutional module, and a channel attention mechanism is used to achieve adaptive recalibration of feature channels. This module establishes channel correlations through global feature statistical modeling and dynamically generates feature channel weight coefficients. This enables the network to autonomously enhance high-value channel responses, including pathological details such as nodule edge features and microcalcifications, while simultaneously mitigating background interference from normal tissue areas. This adaptive feature optimization mechanism effectively improves the model's accuracy in localizing the boundaries of tiny lung nodules and demonstrates enhanced noise suppression capabilities in complex pleural adhesion scenarios.

[0244] 2.3.3 Multi-scale FPN Module

[0245] The task of pulmonary nodule segmentation faces significant multi-scale challenges, as nodules range in size from small nodules in the millimeter range to larger lesions, and their boundaries are often complex and fuzzy. Traditional single-scale feature extraction methods are difficult to simultaneously meet the needs of small target detection and large target segmentation: shallow features retain details but lack semantic depth, while deep features are rich in semantic information but lose boundary details due to reduced resolution. To address this problem, FS-UNet integrates a feature pyramid network (FPN) module between the encoder and decoder, which significantly improves the model's perception and segmentation capabilities for nodules of different sizes through efficient multi-scale feature fusion. As a classic multi-scale architecture, FPN has demonstrated excellent performance in target detection and medical image analysis by fusing high-level semantic information with low-level spatial details, especially in pulmonary nodule segmentation.

[0246] The implementation of FPN in FSUNet relies on three core components: bottom-up path, top-down path and lateral connections. These components together build a systematic structure for multi-scale feature extraction and fusion, such as Figure 7 shown.

[0247] The bottom-up pathway uses the layer-by-layer forward propagation characteristics of the convolutional neural network (CNN) to gradually extract features and compress spatial resolution. This process is achieved through convolution or pooling operations with a stride of 2. The feature maps generated by the encoder are named E1 to E5 from shallow to deep. As the shallowest feature map, E1 retains the fine contours and edge information of the lung structure, such as the boundary details of "Lung", "Right Lung" and "Left Lung". As the deepest feature map, E5 has strong semantic differentiation capabilities and can effectively separate nodules from the background. Feature extraction is achieved through a series of convolution operations. Each layer uses a 3×3 convolution kernel with batch normalization (BN) and ReLU activation function, supplemented by the maximum pooling (MaxPooling) operation to gradually enhance the semantic expression ability of the features, laying the foundation for subsequent multi-scale fusion. Figure 7 As shown in the figure, the bottom-up path starts from the input, and after the "Conv+BN+ReLU" and "MaxPooling" operations, it generates a feature sequence from E1 to E5 layer by layer.

[0248] The top-down pathway aims to propagate high-level semantic information from deep feature maps back to shallow scales and gradually restore spatial resolution through upsampling operations. Starting from E5, the nearest neighbor interpolation method is used to upsample the deep feature maps to align the spatial dimensions with the shallow feature maps (such as E4). This process ensures the effective combination of deep semantic information and shallow spatial details, providing support for subsequent fusion operations. The upsampling operation can be expressed as:

[0249] Ui=Upsample(E(i+1),scale=2) (12)

[0250] Where Ui represents the upsampled feature map, Upsample is the upsampling function, and scale=2 means the resolution is magnified by 2 times. Figure 7 As shown in Figure 1, the top-down path starts from E5 and generates upsampled feature maps U4 to U1 through the green dotted arrows marked as "2xUpsample". These feature maps are aligned with the corresponding Ei feature maps in the spatial dimension, creating conditions for the fusion operation of the horizontal connection.

[0251] Lateral connections fuse features from the bottom-up and top-down pathways to generate multi-scale feature maps, further improving the model's ability to perceive lung nodules of varying sizes. Specifically, for each layer Ei (i = 1, 2, ..., 5), a 1×1 convolution is first used to adjust the number of channels to align with the corresponding upsampled features Ui. Fusion is then performed through element-wise addition. The fusion process can be expressed as:

[0252] Fi'=Ui+Conv 1×1 (Ei)

[0253] Fi'=Concat(Ui,Conv 1×1 (Ei)) (13)

[0254] Among them, Fi' is the fused feature map, Conv1×1 represents the 1×1 convolution operation for channel alignment, and Concat represents the feature concatenation along the channel axis. The fused Fi' contains both deep semantic information and shallow spatial details, significantly improving the expressiveness of the features. These feature maps are finally fed into the FSUNet decoder to generate accurate segmentation predictions. Figure 7 As shown, the lateral connection efficiently integrates the features of Ei and Ui through "1x1 conv" channel adjustments and the "+" operation marked by the red arrow, generating F5' to F1'. F5' is directly derived from E5, while F4' to F1' is achieved through fusion. The decoder receives all fused feature maps Fi' (i = 1, 2, ..., 5) and further integrates the features to produce the final segmentation result.

[0255] 3 Loss Function

[0256] Loss functions are crucial in deep learning models. By quantifying the difference between the predicted value and the true label, they provide guidance for network parameter optimization, thereby reducing errors. In the task of pulmonary nodule segmentation, problems such as sample imbalance, difficulty in segmenting small nodule areas, and blurred boundaries are particularly prominent. To this end, inspired by Yan et al.

[10] , this paper designs a joint loss function that combines DiceLoss

[11] and FocalLoss

[12] to train the FSMedDiff segmentation network based on the diffusion model. This combined loss function effectively balances the accuracy of the overall segmentation area and the ability to focus on difficult samples during training, significantly improving the segmentation accuracy and robustness of the model.

[0257] DiceLoss directly improves the overall accuracy of the segmentation task by optimizing the overlap between the predicted area and the true area. It is particularly suitable for pulmonary nodule segmentation with diverse shapes and fuzzy boundaries. Its mathematical definition is:

[0258]

[0259] Among them, Pi represents the pixel probability predicted by the model, ti is the pixel value of the true label, N is the total number of pixels, ε DL is a smoothing factor to avoid the denominator being zero. By minimizing DiceLoss, the FSMedDiff model can effectively capture the morphological characteristics of lung nodules and significantly improve the segmentation consistency of large-sized nodules and areas with blurred boundaries.

[0260] To address the problems of sample imbalance and difficulty in segmenting small nodules, this paper introduces FocalLoss, which dynamically adjusts the weights of difficult-to-classify samples to enhance the model's ability to focus on small lesions and low-contrast areas. The formula is:

[0261] FocalLoss=-α FL (1-p t ) γ log(p t ) (15)

[0262] Among them, pt is the predicted probability of the target category, α FL is a balancing factor (set to 0.25 in this paper), and γ is a tuning factor (set to 2) used to control the weighting of difficult-to-classify samples. FocalLoss reduces the loss contribution of easy-to-classify samples, allowing the model to focus on small nodules and boundary areas that are difficult to segment.

[0263] The joint loss function designed in this invention integrates the advantages of DiceLoss and FocalLoss, and its form is:

[0264] L loss =DiceLoss+τ·FocalLoss (16)

[0265] Here, τ = 1 is a weighting factor, which has been experimentally verified to effectively balance the contributions of the two losses. DiceLoss ensures the accuracy of the overall segmentation region, while FocalLoss increases the focus on difficult samples. The synergy of the two significantly enhances the performance of the FSMedDiff model.

[0266] In order to evaluate the effect of the joint loss function in pulmonary nodule segmentation, the present invention conducted an in-depth analysis of the loss change trend during the training process. The results are as follows: Figure 8As shown in the figure. Early in training, the loss value dropped rapidly from approximately 1.2 to around 0.3, indicating that the model was able to quickly learn the basic patterns of the training data in the early stages. However, during epochs 50-150, the loss fluctuated significantly between 0.2 and 0.4. This may be due to two factors: first, the instability of model parameter updates, and second, the diversity of lung nodule data, including differences in nodule size, shape, location, and contrast, which posed a challenge to the optimization process. As training progressed, the loss fluctuations gradually decreased and eventually stabilized at around 0.15, indicating that the model overcame the instability of early training through continuous optimization and achieved stable convergence. Figure 8 The key role of the joint loss function in improving model robustness and segmentation accuracy was further verified, especially when dealing with small lesions and areas with blurred boundaries, which effectively alleviated the sample imbalance problem and significantly improved segmentation accuracy and stability.

[0267] 4 Experiments

[0268] Dataset

[0269] This paper uses the LIDC-IDRI dataset

[13] for the lung nodule segmentation task. The dataset contains 6474 chest images, which are divided into training set and test set in a ratio of 9:1.

[0270] Experimental parameter settings

[0271] The experimental environment was Ubuntu 20.04, and the hardware configuration included an Intel Xeon Platinum 8276 processor and a GeForce RTX 3090 GPU with 24GB of video memory. Stochastic Gradient Descent (SGD) was used for optimization, with an initial learning rate of 1e-5, dynamically adjusted to accommodate different training phases using the Annealing Scheduler. The weight decay coefficient was 0.0005, the batch size was 32, and the total number of model iterations was 300. All input images were preprocessed and resized to 256×256.

[0272] Evaluation indicators

[0273] In order to comprehensively evaluate the performance of the FSMedDiff model in the lung nodule segmentation task, this paper selected the following key indicators: Dice similarity coefficient (DSC), precision, average precision (AP), mean intersection over union (mIoU), and frames per second (FPS). These indicators quantitatively analyze the model from three dimensions: segmentation accuracy, detection capability, and inference efficiency. The following is a detailed definition and explanation of these evaluation indicators:

[0274] DSC

[14] , also known as the Dice similarity coefficient, is a commonly used metric to measure the similarity between two sets, and is particularly widely used in medical image segmentation. It evaluates the accuracy of segmentation by calculating the degree of overlap between the predicted result and the true label. Its calculation formula is:

[0275]

[0276] Where A represents the predicted segmentation area, B represents the ground-truth label area, |A∩B| represents the number of pixels in the area where the prediction overlaps with the ground-truth label, and |A| and |B| represent the number of pixels in the predicted area and the ground-truth label area, respectively. The value of DSC ranges between 0 and 1. The closer the value is to 1, the closer the segmentation result is to the ground-truth label, and the higher the accuracy.

[0277] mIoU

[15] evaluates the multi-category segmentation capability of the model by calculating the intersection over union (IoU) of the predicted area and the true area. The formula is:

[0278]

[0279] Where K is the number of categories (typically 2 in the lung nodule task, i.e., nodules and background), Sp,k and Sg,k are the predicted and true regions of category k, respectively. A higher mIoU value indicates a stronger ability of the model to distinguish between nodules and background, making it suitable for evaluating segmentation consistency in complex backgrounds.

[0280] FPS

[16] measures the inference efficiency of the model and represents the number of image frames processed per second. Its calculation formula is:

[0281]

[0282] Where N is the number of images processed and T is the time required for running.

[0283]

[0284] AP

[17] evaluates the detection performance of the model at different confidence thresholds by calculating the area under the precision-recall curve. The formula is:

[0285] Here, P(R) represents the precision under different recall rates R. In pulmonary nodule segmentation, AP is particularly suitable for evaluating the model's ability to identify small, highly complex nodules, especially in the case of sample imbalance (such as a low proportion of nodules), and can comprehensively reflect the robustness of the model.

[0286] 4.4 Experimental Results and Analysis

[0287] This section systematically evaluates the performance of the FSMedDiff model in the pulmonary nodule segmentation task through comparative experiments, ablation experiments, and parameter analysis, and deeply explores the contribution of each module and its hyperparameters to the model effect.

[0288] 4.4.1 Analysis of comparative experimental results

[0289] To verify the effectiveness of FSMedDiff, we compared its performance with multiple SOTA segmentation models on the LIDC-IDRI dataset. The quantitative results are shown in Table 4.1.

[0290] Table 4.1 Performance comparison with mainstream methods

[0291]

[0292]

[0293] FSMedDiff achieved 0.915 on DSC, significantly outperforming ResNet-50 (0.651) and U-Net++ (0.728), indicating that it can more accurately capture the boundaries and morphological features of lung nodules. This is due to its innovative multi-scale feature fusion design and anatomical prior embedding. Compared with traditional convolutional networks (such as ResNet-50) or deep segmentation models (such as U-Net++), FSMedDiff shows higher adaptability in complex lung nodule segmentation. In terms of Precision, FSMedDiff scored 0.973, far exceeding U-Net (0.825) and Mask R-CNN (0.837), reflecting its superiority in reducing missegmentation, especially for clinical scenarios with strict requirements for high accuracy. The AP score was 0.907, slightly higher than LSegDiff (0.904), indicating that its comprehensive detection capabilities under different confidence thresholds are at the leading level, especially for small or low-contrast nodules. On mIoU, FSMedDiff reaches 0.953, second only to DiffUNet (0.984), further verifying its competitiveness in pixel-level segmentation accuracy.

[0294] From the FPS point of view, FSMedDiff's FPS is 35.5, which is better than Mask R-CNN (8.6) and U-Net (27.9), but lower than DiffUNet (68.4). Although it is not as fast as some lightweight models, FSMedDiff has achieved a balance between accuracy and efficiency. Figure 9 As shown in the figure, it is at the Pareto frontier in the accuracy-speed trade-off diagram, indicating that it meets the real-time requirements without significantly sacrificing the segmentation performance, and is particularly suitable for lung nodule segmentation tasks in complex backgrounds.

[0295] 4.4.2 Analysis of Ablation Experiment Results

[0296] To further validate the effectiveness and contribution of each key component in the FSMedDiff model, we designed a series of ablation experiments to systematically evaluate the impact of the Anatomy Controller (AC), the Attention Module (SENet, SE), and the Multi-Scale Feature Fusion Module (FPN) on lung nodule segmentation performance. The experiments incrementally added each module to the base model to observe its impact on segmentation accuracy. All experiments maintained consistent network parameters and settings to ensure comparability of the results. The experiments were conducted on the LIDC-IDRI public dataset, and the results are shown in Table 4.2.

[0297] Table 4.2 Performance comparison on the LIDC-IDRI dataset

[0298]

[0299] As shown in the results in Table 4.2, when no modules are used, the model's segmentation performance is poor (DSC = 0.344, Precision = 0.547, mIoU = 0.759), indicating that the base network, without targeted optimization, struggles to effectively capture the complex features of lung nodules. When the AC module is introduced alone, the DSC improves to 0.806, the Precision reaches 0.853, and the mIoU is 0.903, demonstrating that the AC module significantly enhances the model's feature representation capabilities by providing anatomical prior information. However, its sole effect is still insufficient to achieve high-precision segmentation, likely due to a lack of sufficient attention to detailed features. In contrast, when the SE module is used alone, the DSC is 0.796, the Precision is 0.845, and the mIoU is 0.864, indicating that the attention mechanism can effectively filter key features and improve segmentation accuracy, but the improvement is smaller than that of the AC module. When the FPN module is added alone, the DSC is 0.654, the Precision is 0.802, and the mIoU is 0.885, showing its role in multi-scale feature fusion. However, the effect is limited when used alone, which may be limited by the lack of support from prior information or feature selection capabilities.

[0300] Further analysis of the module combination effect shows that when AC and SE are combined, DSC increases to 0.869 and Precision reaches 0.885, indicating a synergistic effect between the two, integrating anatomical information with key feature selection to optimize segmentation results. When SE and FPN are combined, DSC reaches 0.801 and mIoU significantly increases to 0.915, demonstrating the complementarity of the attention mechanism and multi-scale fusion in improving the intersection-over-union ratio. When AC, SE, and FPN are fully combined, DSC reaches 0.915, Precision reaches 0.947, and mIoU reaches 0.895, achieving optimal performance in all indicators. This result demonstrates that the three modules working together can fully leverage their respective strengths, forming a powerful feature extraction and fusion capability, and achieving high-precision lung nodule segmentation.

[0301] 4.4.3 Parameter performance results analysis

[0302] To systematically evaluate the impact of the weight parameter λ in the Anatomy Controller module on lung nodule segmentation performance (see Section 4.3.4 for details), we conducted ablation experiments based on the LIDC-IDRI dataset. While maintaining the same experimental conditions, we set λ to 0.1, 0.3, 0.5, 0.7, and 0.9 for comparative analysis. The experimental results are shown in Table 4.3.

[0303] Table 4.3 Comparison of model performance under different λ values

[0304]

[0305] As shown in Table 4.3, when λ = 0.5, the model's DSC and AP reach optimal values ​​of 0.869 and 0.853, respectively. Analysis shows that smaller λ values ​​(such as 0.1 and 0.3) result in insufficient regularity in the latent space, making it difficult for the model to effectively capture the key features of lung nodules, resulting in decreased performance. Conversely, larger λ values ​​(such as 0.7 and 0.9), while enhancing the regularity of the latent space, overly constrain the model's expressive power, reducing reconstruction quality and segmentation accuracy.

[0306] 5. Visual Analysis

[0307] Figure 10 This example demonstrates the segmentation performance of the FSMedDiff model on the LIDC-IDRI dataset. Comparing the output of the second column, "Input," with the third column, "FSMedDiff," demonstrates that the model accurately identifies lung nodules and lesions, clearly delineating target boundaries and achieving high-precision segmentation. This model demonstrates high accuracy and excellent generalization, helping physicians quickly locate nodules and thereby improving the efficiency and quality of clinical diagnosis.

[0308] 1. Specific application fields or related products of the present invention.

[0309] 1. Hospital imaging diagnostic auxiliary system:

[0310] This invention can be directly embedded into a hospital's lung CT film reading system. After a doctor uploads a lung image, the system automatically segments the lung nodule area and generates a clear contour image and relevant indicators (such as nodule volume and morphological parameters). This helps doctors quickly determine whether the nodule is benign or malignant, improving the efficiency of early lung cancer screening. This can effectively reduce the burden on doctors and improve the efficiency and accuracy of film reading, especially in daily batch film reading and high-intensity work scenarios.

[0311] 2. Medical imaging teaching and training system:

[0312] It can be integrated into medical school or hospital training systems as an auxiliary teaching tool. By observing the automatic segmentation process of real lung nodule CT images, medical students and young doctors can understand typical lesion characteristics and segmentation algorithm mechanisms, which helps improve image interpretation skills, shorten the learning curve, and promote the acceptance and use of AI-assisted tools among doctors.

[0313] 2. Relevant evidence of the technical effects obtained by the embodiments of the present invention.

[0314] like Figure 11 Precision, AP, and mIoU comparison chart: The three indicators are displayed side by side, intuitively reflecting the comprehensive advantages of FSMedDiff in accuracy, average precision, and average intersection over union.

[0315] like Figure 12 DSC comparison chart: shows the Dice similarity coefficient of each method on LIDC-IDRI, with FSMedDiff reaching the highest 0.915.

[0316] like Figure 13 Inference Speed ​​Comparison Chart: Comparison of the FPS performance of each model. DiffUNet leads in speed, while FSMedDiff achieves a balance between speed and accuracy.

[0317] 1. Significantly improve segmentation accuracy

[0318] In tests on the LIDC-IDRI public dataset, the proposed method achieved a Dice Similarity Coefficient (DSC) of 0.915, significantly outperforming ResNet-50 (0.651) and U-Net++ (0.728). A DSC close to 1 indicates that the model's predictions are highly consistent with expert annotations, demonstrating that the proposed method is more accurate in nodule boundary identification and morphological restoration.

[0319] 2. Possess real-time reasoning capabilities, suitable for clinical workflows

[0320] This method achieves an inference speed of 35.5 frames per second (FPS), exceeding U-Net (27.9 FPS) and Mask R-CNN (8.6 FPS), meeting the batch processing needs of hospitals. Although slightly slower than DiffUNet, this method strikes a good balance between speed and accuracy, reaching the Pareto frontier on the "accuracy-speed" graph and demonstrating strong practicality.

[0321] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0322] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.

Claims

1. A lung nodule segmentation method based on a diffusion model, characterized in that: The following steps are involved: Step 1: A latent space fusion method based on anatomical constraints is used to integrate anatomical prior knowledge such as lung tissue contours through a variational autoencoder and combined with a KL divergence constraint strategy; Step 2: Build a multimodal collaborative optimization FSUNet architecture, using a cross-scale feature pyramid fusion mechanism, combined with channel attention gating and dynamic noise scheduling algorithms; Step 3: Design the Dice-Focal joint loss function to suppress complex background noise interference, and at the same time generate noise samples based on the forward diffusion strategy.

2. The pulmonary nodule segmentation method based on the diffusion model according to claim 1, characterized in that: The forward diffusion: The forward diffusion process simulates image degradation and gradually adds Gaussian noise to the initial latent variable zp (usually generated by the input image encoding) through a Markov chain, generating a series of noised states until zT; the process can be formally expressed as: in, is the cumulative noise scheduling parameter, α s =1-β s ,β s ∈(0,1) is the predefined noise intensity, t represents the time step, ∈ t is the random noise sampled from the standard normal distribution N(0,I); the noise intensity β s Usually a preset linear or cosine scheduling strategy is adopted; Backward denoising is the inverse process of forward diffusion, and its goal is to gradually recover the original latent variables from the highly noisy state zT FSUNet, as a denoising network, receives the current diffusion state zt and time step information t as input and predicts noise based on the deep learning model. The denoising process updates the latent variables by the following formula: Among them, σ t is the posterior variance, usually set to σ t =β t , η is random noise, which is used to enhance sampling diversity; FSUNet extracts multi-scale features through the encoder, combines with SENet to dynamically adjust channel weights to focus on key information, and uses FPN to fuse shallow details with deep semantics to optimize noise prediction accuracy; the denoising process iterates from t = T to t = 0, and finally generates the denoised latent variable Decoding and segmentation mask generation Denoised latent variables The input pre-trained VAE decoder Ds is mapped to a feature representation with the same spatial dimensions as the input image through multi-layer deconvolution operations; the decoding process can be expressed as: in, Based on the recovered image latent representation, a lung nodule segmentation mask Mnodule is subsequently generated through thresholding or post-processing steps.

3. The pulmonary nodule segmentation method based on the diffusion model according to claim 1, characterized in that: The suppression of complex background noise interference: Anatomy Controller Module 1) Input processing and anatomical supervision The Anatomy Controller module takes the original lung nodule image x as input and introduces the anatomical mask ma as a supervisory signal. The anatomical mask ma contains information such as lung contours, tracheal position, and key vascular structures, aiming to enhance the model's understanding and learning of lung anatomical features. 2) Encoding stage: latent representation generation In the encoding stage, the module uses the pre-trained VAE encoder Ev to extract features from the input image x; through multi-layer convolution operations and nonlinear activation functions, the encoder gradually compresses the high-dimensional image information into the parameters of the potential distribution, that is, the mean μ a and log variance The process can be expressed as: Subsequently, the latent variable za follows a multivariate normal distribution To ensure the differentiability of the training process, the reparameterization technique is used to calculate za: With a =μ a +∈ a ⊙σ a (5) Among them, ∈a is random noise sampled from the standard normal distribution, and ⊙ represents element-by-element multiplication. This process not only preserves the semantic information of the image; 3) Decoding stage: latent space mapping and noise injection z p =D v (z a )+F(z a )(6) In the decoding stage, the latent variable za is input to the VAE decoder Dv; the decoder maps the high-dimensional latent representation za to a low-dimensional latent variable zp through multiple layers of deconvolution and upsampling operations, providing input for the subsequent FSUNet denoising network; in this process, the supervisory signal of the anatomical mask ma is embedded in the training target of the decoder to guide the model to learn the relevant features of the anatomical structure; the latent space mapping process can be formalized as: Where F(·) is a nonlinear mapping function used to adjust the uniformity and expressiveness of the potential space distribution. To further enhance the robustness of the potential representation and adapt to complex segmentation scenarios, random Gaussian noise is introduced in the mapping process: Here, σ n is the noise intensity hyperparameter; noise injection not only enriches the diversity of the latent space, but also simulates the noise interference that may exist in real lung images, thereby improving the generalization ability of the model; 4) Loss function optimization Design a comprehensive loss function including reconstruction loss and KL divergence loss; reconstruction loss Lrecon is used to measure the decoder reconstructed image The similarity between the original image x is defined using the mean square error (MSE): KL divergence loss LKL constrains the latent variable distribution q(z a |x) is close to the prior distribution N(0,I) to improve the regularity and generalization ability of the latent space: Where d is the dimension of the latent space; the comprehensive loss function is: A_Loss=L recon +λ·L KL (10) The weight parameter λ balances the relationship between reconstruction accuracy and latent space regularity. Through multiple rounds of experimental evaluation, it was found that λ = 0.5 can achieve better performance between reconstruction quality and model stability.

4. The pulmonary nodule segmentation method based on the diffusion model according to claim 1, characterized in that: The FSUNet architecture: 1) VAE-guided information fusion The pre-trained Es provides prior knowledge of the lung anatomical structure by learning the potential distribution of the image, effectively guiding the denoising and restoration process. Specifically, the latent variable zs output by Es is fused with the intermediate feature map of the encoder through a skip connection, which can be mathematically expressed as: Among them, F enc is the feature map extracted by the encoder, represents the feature fusion operation, z s ∈R d is the latent variable, and d is the latent space dimension. This process not only enhances the model’s sensitivity to lung nodule details, but also improves the stability of the denoising process. 2) Adaptive SE module FSUNet enhances its ability to capture key features of lung nodules by embedding the SENet attention mechanism. In the two-stage encoder-decoder architecture, the network cascades SENet units after each 3×3 convolutional module, achieving adaptive recalibration of feature channels through the channel attention mechanism. 3) Multi-scale FPN module; 4) Loss function.

5. The lung nodule segmentation method based on the diffusion model according to claim 4, characterized in that: The multi-scale FPN module: The top-down path aims to transfer high semantic information in deep feature maps back to shallow scales and gradually restore spatial resolution through upsampling operations. Starting from E5, the nearest neighbor interpolation method is used to upsample the deep feature maps to align with the spatial size of the shallow feature maps. This process ensures the effective combination of deep semantic information and shallow spatial details, providing support for subsequent fusion operations; the upsampling operation can be expressed as: Ui=Upsample(E(i+1),scale=2) (12) Where Ui represents the upsampled feature map, Upsample is the upsampling function, and scale = 2 means the resolution is magnified by 2 times. The top-down path starts from E5 and generates the upsampled feature maps U4 to U1 through the green dotted arrow marked "2xUpsample". These feature maps are aligned with the corresponding Ei feature maps in the spatial dimension, creating conditions for the fusion operation of the horizontal connection; The lateral connection generates a multi-scale feature map by fusing the features of the bottom-up and top-down pathways, further improving the model's ability to perceive lung nodules of different sizes. Specifically, for each layer Ei (i = 1, 2, ..., 5), the number of channels is first adjusted using 1×1 convolution to keep it consistent with the corresponding upsampled features Ui. Subsequently, the fusion is completed through element-wise addition. The fusion process can be expressed as: Fi'=Ui+Conv 1×1 (They) Fi'=Concat(Ui,Conv 1×1 (They)) (13) Among them, Fi' is the fused feature map, Conv1×1 represents the 1×1 convolution operation for channel alignment, and Concat represents the feature splicing along the channel axis; the fused Fi' contains both deep semantic information and shallow spatial details, significantly improving the expressiveness of the features; these feature maps are finally fed into the FSUNet decoder to generate accurate segmentation predictions.

6. The lung nodule segmentation method based on the diffusion model according to claim 4, characterized in that: The loss function: The loss function is crucial in deep learning models by quantifying the difference between the predicted value and the true label; its mathematical definition is: Among them, Pi represents the pixel probability predicted by the model, ti is the pixel value of the true label, N is the total number of pixels, ε DL is a smoothing factor to avoid zero denominator; This paper introduces FocalLoss to enhance the model's ability to focus on small lesions and low-contrast areas by dynamically adjusting the weights of difficult-to-classify samples. The formula is: FocalLoss=-a FL (1-p t ) γ log(p t ) (15) Among them, pt is the predicted probability of the target category, α FL is a balancing factor (set to 0.25 in this paper), and γ is an adjustment factor (set to 2) used to control the weighting degree of difficult-to-classify samples; FocalLoss reduces the loss contribution of easy-to-classify samples, allowing the model to focus on small nodules and boundary areas that are difficult to segment; The designed joint loss function integrates the advantages of DiceLoss and FocalLoss, and its form is: L loss =DiceLoss+τ·FocalLoss (16) Among them, τ=1 is the weight factor, which has been experimentally verified to effectively balance the contributions of the two losses; DiceLoss ensures the accuracy of the overall segmentation area, while FocalLoss improves the attention to difficult samples. The synergistic effect of the two significantly enhances the performance of the FSMedDiff model.

7. A pulmonary nodule segmentation system based on a diffusion model that implements the pulmonary nodule segmentation method based on a diffusion model according to any one of claims 1 to 6, characterized in that: The pulmonary nodule segmentation system based on the diffusion model includes: The fusion module is used to adopt an anatomically constrained latent space fusion method, integrating anatomical prior knowledge such as lung tissue contours through a variational autoencoder and combining it with a KL divergence constraint strategy; A building block for constructing the FSUNet architecture for multimodal collaborative optimization, utilizing a cross-scale feature pyramid fusion mechanism, coupled with channel attention gating and dynamic noise scheduling algorithms; The suppression module is used to design the Dice-Focal joint loss function to suppress complex background noise interference, while also generating noise samples based on the forward diffusion strategy.

8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the lung nodule segmentation method based on the diffusion model as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the pulmonary nodule segmentation method based on the diffusion model according to any one of claims 1 to 6.

10. An information data processing terminal, characterized in that: The information data processing terminal is used to implement the lung nodule segmentation system based on the diffusion model as described in claim 7.

Citation Information

Cited By

  • Lung ventilation-perfusion development area image segmentation and quantitative analysis system

    CN120953301A

  • Medical image X-ray pulmonary tuberculosis medical information management system based on AI

    CN121354936A

  • PET / CT head and neck tumor automatic segmentation method based on fusion diffusion model

    CN121437878A

  • Medical image segmentation method and system based on diffusion difference learning

    CN121811412A

  • Pulmonary nodule MRI segmentation method based on attention-guided cross-modal fusion

    CN121904091A