Medical image denoising and segmentation integrated model trained based on conductible diffusion method

By cascading training of the improved denoising diffusion module and the joint motion and segmentation module, the problem of the inability to coordinate the denoising and segmentation tasks in existing medical image denoising models is solved, and high-precision and high-robustness medical image segmentation is achieved.

CN121304615APending Publication Date: 2026-01-09BEIJING LUHE HOSPITAL AFFILIATED TO CAPITAL MEDICAL UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511488602.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

Existing medical image denoising models cannot achieve synergistic optimization of denoising and segmentation tasks, resulting in damage to image details and reduced segmentation performance during the denoising process.

Method used

A medical image denoising and segmentation integrated model trained based on the differentiable diffusion method is adopted. The improved denoising diffusion module and the joint motion and segmentation module are cascaded for training. The joint loss function is used to realize gradient backpropagation and parameter synchronous optimization. Combined with J-invariance theory, Markov chain state matching and time-coded U-Net denoising unit, end-to-end collaborative optimization of denoising and segmentation is achieved.

Benefits of technology

It improves the accuracy and robustness of medical image segmentation, ensures the physiological rationality of segmentation results, solves the problems of image detail loss and low segmentation quality caused by the non-differentiability of traditional denoising methods, and realizes high-quality integrated denoising and segmentation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121304615A_ABST
    Figure CN121304615A_ABST
Patent Text Reader

Abstract

The invention provides a medical image denoising and segmentation integrated model trained based on a conductible diffusion method, and belongs to the field of medical image processing, and the model comprises an improved denoising and diffusion module which is used for receiving an initial medical image, predicting an original clean signal of the initial medical image based on an improved denoising and diffusion probability model, and obtaining a final denoised medical image; the joint motion and segmentation module is used for receiving the initial medical image, extracting features through a shared encoder, and performing joint learning through a motion estimation branch and a segmentation branch to obtain segmentation data; the cascade training mechanism carries out end-to-end training on the improved de-noising diffusion module and the joint motion and segmentation module through a derivable connection, gradient back propagation is realized by using a joint loss function, and cascade training is completed. According to the method, the problem that the quality of segmented medical images is reduced due to the fact that denoising and segmentation tasks cannot be collaboratively optimized due to the fact that a traditional denoising method cannot be guided in sampling and is difficult to carry out cascade training with a segmentation model is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing, and in particular relates to an integrated medical image denoising and segmentation model trained based on the guided diffusion method. Background Technology

[0002] In the field of medical imaging, accurate segmentation of medical images has a significant impact on the subsequent assessment of functional indicators of the affected area. However, noise is almost unavoidable in current medical imaging examinations, which poses a challenge to accurate segmentation of medical images.

[0003] Current mainstream denoising methods are based on deep learning. The Noise2Noise model proposes unsupervised training, using noisy images for training, and improves image segmentation and training strategies, proposing models such as the 4DMR image denoising model. The denoising diffusion probability model is a generative model with two processes: forward diffusion and backward denoising. Compared with Generative Adversarial Networks (GANs), it has better convergence and generates higher-quality samples. In recent years, it has been frequently used for medical image denoising, such as retinal image denoising and multimodal image denoising. However, while current mainstream denoising models can maximize noise reduction and restore anatomical structures, the denoising sampling process cannot be differentiated and cannot be cascaded with the segmentation model. This results in a lack of synergistic optimization between the denoising and segmentation tasks, limiting overall performance improvement, leading to loss of image details, blurred edges, and even impaired segmentation performance. Summary of the Invention

[0004] To address the aforementioned shortcomings in existing technologies, the medical image denoising and segmentation integrated model based on the differentiable diffusion method provided by this invention solves the problem that traditional denoising methods, due to limitations of the denoising model and the non-differentiable sampling process, are difficult to cascade with the segmentation model, resulting in the inability to coordinate and optimize denoising and segmentation tasks, ultimately leading to a decline in medical image quality.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A medical image denoising and segmentation integrated model trained based on the differentiable diffusion method includes: An improved denoising diffusion module is used to receive the initial medical image, predict the original clean signal of the initial medical image based on the improved denoising diffusion probability model, and obtain the final denoised medical image. The joint motion and segmentation module receives initial medical images, extracts features through a shared encoder, and obtains segmentation data through joint learning via motion estimation and segmentation branches. Cascaded training mechanism: The improved denoising diffusion module and the joint motion and segmentation module are trained end-to-end through a differentiable connection. The joint loss function is used to realize gradient backpropagation and parameter synchronous optimization to complete the cascaded training. The segmentation data obtained by the joint motion and segmentation module after cascaded training is used as the denoised segmented medical image.

[0006] To address the issues of low-quality medical image segmentation caused by traditional denoising methods that overemphasize denoising and thus damage key structural details of images, and the inability of existing denoising diffusion models to be jointly optimized with segmentation models due to non-differentiable sampling, this invention integrates an improved, differentiable denoising diffusion module with a joint motion estimation segmentation module to construct a cascaded model that can be trained end-to-end. By utilizing a joint loss function to achieve gradient backpropagation and collaborative optimization, the contradiction between denoising and segmentation objectives is resolved, improving the accuracy and robustness of medical image segmentation under complex noise conditions. Simultaneously, the motion estimation branch ensures the physiological rationality of the segmentation results, providing reliable technical support for the accurate analysis of medical images.

[0007] Furthermore: the improved noise reduction and diffusion module includes: The preliminary image denoising N2N unit is used to perform preliminary denoising on the initial medical image based on the J-invariance theory to obtain the initial denoised medical image. The noise level estimation unit is used to calibrate the residual noise based on the residual distribution between the initial denoised medical image and the initial medical image. Based on the residual noise, the minimum timestamp of the diffusion process is calculated using the Markov chain state matching method. The U-Net denoising unit with time coding is used to predict the original clean signal of the initial medical image by using the minimum timestamp as the time step in a forward propagation manner, so as to obtain the final denoised medical image.

[0008] The further beneficial effects mentioned above are as follows: the preliminary image denoising N2N unit achieves unsupervised preliminary denoising based on J-invariance theory, reducing image noise; the noise level estimation unit adaptively determines the minimum diffusion timestamp through Markov chain state matching, determining the time step from the initial noisy medical image to start denoising, rather than the conventional time step from the noise itself; the time-coded U-Net model achieves differentiability of the denoising process by predicting the original clean signal instead of noise; the improved denoising diffusion module can effectively suppress image noise while preserving structural details crucial to the segmentation task, ensuring gradient connectivity with the subsequent joint motion and segmentation modules, providing a technical basis for end-to-end cascaded training, and improving the segmentation quality of medical images in real clinical noise environments.

[0009] Furthermore, the loss function expression for the preliminary image denoising N2N unit is as follows:

[0010]

[0011]

[0012] in, The loss function for the initial image denoising N2N unit is... For initial denoised medical images, For norm, Denoising function Through aggregation -1 volumetric slices are used as the input image. For initial medical images containing noise, For clean medical imaging, For constant terms, and All are linear coefficients. Additive noise, With a mean of 0 and a covariance matrix of The normal distribution This is a unit matrix consistent with the dimensions of medical imaging.

[0013] The further beneficial effects mentioned above are as follows: This invention combines J-invariance theory with the Noise2Noise model geometrically, and utilizes the noisy image itself for unsupervised learning, which can obtain clean medical images. At the same time, the design of the loss function enables the initial image denoising N2N unit to directly learn the denoising mapping from the noisy data. By minimizing the difference between noisy image pairs, the denoising effect is improved, the stability of the training process is maintained, and high-quality initial denoising results are provided for subsequent diffusion denoising, ensuring the high quality of subsequent segmentation tasks.

[0014] Furthermore, the expression for the calibration residual noise is as follows:

[0015]

[0016]

[0017]

[0018] in, For residual noise, For initial medical images containing noise, Clean medical images for initial image denoising N2N unit loss function estimation. This represents the mean of the residual noise. The adjusted residual noise, For the adjusted denoised medical images, and All are balance coefficients.

[0019] The further beneficial effects mentioned above are: the present invention first calculates the residual noise. Quantize residual noise, then use the mean. Calibration residuals are used to match the time steps of the noise adaptation. By calibrating residual noise, adaptive denoising is achieved, avoiding excessive smoothing of noise and loss of details. It also makes the cascaded denoising and segmentation differentiable, solving the problems of residual noise shift and model non-differentiability caused by the fixed time step in traditional denoising diffusion models.

[0020] Furthermore, the distance expression for the matching states of the Markov chain is as follows:

[0021] ,

[0022] in, Minimum timestamp For a predefined noise schedule, Input image and predefined noise timetable All possible posterior standard deviations For hyperparameters, Let be the probability density function. The initial medical image dimensions are input. It is a natural exponential function. For the residual noise after calibration, To fit the calibrated residual noise to a value with a mean of 0 and a covariance matrix of... Multidimensional Gaussian distribution .

[0023] The further beneficial effect mentioned above is that the matching state of the Markov chain determines the position of the input medical image in the Markov chain, so that subsequent backsampling can start from this state to generate denoised medical images, thereby improving the denoising quality and denoising efficiency.

[0024] Furthermore, the expression for the final denoised medical image is as follows:

[0025]

[0026]

[0027] in, This is the final denoised medical image obtained after the reverse process of the entire diffusion model. For all possible values ​​of the desired clean, noise-free medical image, This represents the joint probability distribution of medical images during the inverse denoising process. This represents the total time step of the diffusion process. The minimum timestamp is the time step of the current diffusion process. For time step Noisy medical images, The initial noise distribution at the start of the diffusion process. It serves as the core learning component for neural network parameterization of single-step denoising transition probabilities. For mean prediction, For variance prediction, These are the parameters of the neural network trained on a medical image dataset.

[0028] The further beneficial effect mentioned above is: joint probability distribution With conditional normal distribution This makes the denoising process differentiable, supports end-to-end cascaded training of denoising and segmentation, avoids the accumulation of errors from separate training, and solves the problems of non-differentiable sampling, difficulty in cascading denoising and segmentation, and insufficient denoising accuracy in traditional denoising diffusion models, thus providing high-quality denoised medical images.

[0029] Furthermore: the joint motion and segmentation module specifically includes: A shared encoder is used to perform multi-scale feature extraction and manual annotation on initial medical images to obtain pre-organ movement temporal images, manually annotated pre-organ movement temporal images, post-organ movement temporal images, and manually annotated post-organ movement temporal images. The motion unit is used to stitch together manually annotated images of organ activity before and after movement to predict the deformation field of the activity process between the first and second phase images. The segmentation unit is used to segment manually annotated images of organ activity before and after movement using a fully convolutional network, obtaining segmentation results for the images before and after organ activity.

[0030] The further beneficial effects mentioned above are as follows: In response to the problems of large segmentation errors and feature loss caused by the failure of traditional medical image segmentation to utilize organ motion temporal information and to consider the processing of organ motion and segmentation, this invention extracts multi-scale features through a shared encoder, taking into account both the details and global information of medical images, and obtains the deformation field of organ motion process through segmentation units, thereby improving the accuracy of medical image segmentation and solving the problem of low medical image segmentation quality caused by organ periodic motion.

[0031] Furthermore, the expression for the joint loss function is as follows:

[0032]

[0033]

[0034]

[0035]

[0036]

[0037]

[0038] in, The loss function for the joint motion and segmentation module, The weights of the loss function are used to estimate the motion. For motion estimation loss function, To divide the weights of the loss function, For the segmentation loss function, For the improved loss function of the denoising diffusion module, For the improved noise reduction and diffusion module, For segmentation modules, Images of organ activity prior to movement. These are phase images of organ activity following movement. Manual annotation of pre-movement phase images of organ activity. Manual annotation of pre-movement phase images of organ activity was obtained. The segmentation results are for the temporal images before organ activity and movement. The segmentation results are for the temporal images after organ activity and movement. These are temporal images of organ activity before movement, after denoising processing. These are temporal images of organ activity after denoising processing. for The spatial x-coordinate of a sampled pixel in an image. This represents the spatial ordinate of the pixel in the image. To predict the displacement in the lateral direction of the deformation field Φ during the activity process between the pre-movement phase image and the post-movement phase image of organ activity, To predict the displacement in the longitudinal direction of the deformation field Φ during the motion process between the pre-movement phase image and the post-movement phase image of organ activity, It is the L2 norm. Here is the Huber loss function.

[0039] The further beneficial effects mentioned above are as follows: The present invention designs a composite loss function to fuse the motion unit and the segmentation unit for training, ensuring that the deformation field representing the organ motion process conforms to the real physiological law, while improving the image segmentation accuracy, so that the joint motion and segmentation module can learn the accurate deformation field and segmentation results at the same time, realize the synergistic optimization of organ motion phase and segmentation results, and improve the segmentation performance of the joint motion and segmentation module.

[0040] Furthermore, the cascaded training mechanism specifically includes: Through a differentiable gradient path, the loss gradient of the joint motion and segmentation module is backpropagated to the improved denoising and diffusion module. The parameters of the improved denoising and diffusion module and the parameters of the joint motion and segmentation module are jointly supervised through the joint loss function. Cascaded training is completed according to the joint loss function. The joint motion and segmentation module after cascaded training is used to directly output segmentation data with reduced noise interference based on the initial medical images, resulting in the final denoised medical images.

[0041] The further beneficial effects mentioned above are as follows: by establishing a differentiable gradient path, the improved denoising diffusion module and the segmentation module are trained together, enabling denoising and segmentation to be trained together, thereby improving segmentation quality and efficiency; through joint supervision of the joint loss function, the joint motion and segmentation module can directly output segmentation data with noise interference removed from the original medical images, improving the segmentation accuracy and robustness of the model under noise interference, and ensuring the reliability of the output results.

[0042] The beneficial effects of this invention are: The improved denoising diffusion model of this invention predicts the original clean signal instead of noise and combines Markov chain state matching to adaptively calculate the minimum timestamp. This solves the problems of traditional filtering and denoising losing boundary details in medical images and ordinary denoising diffusion models causing over-denoising or incomplete denoising due to fixed time steps. It ensures that the final denoised medical image still retains the fine structure of organs and provides a high-quality image foundation for subsequent segmentation tasks. The joint motion and segmentation module extracts multi-scale features through a shared encoder, and motion estimation utilizes the temporal phase of organ periodic motion. The segmentation branch optimizes the segmentation boundary using a fully convolutional network. The two branches work together to solve the problem that traditional segmentation does not utilize the temporal information of organ motion and does not consider the processing of organ motion and segmentation, resulting in large image segmentation errors. Therefore, the improved denoising diffusion model and the joint motion and segmentation module can be cascaded to achieve end-to-end collaborative optimization of denoising and segmentation, avoiding the error propagation of separate training. Attached Figure Description

[0043] Figure 1This is a structural diagram of a medical image denoising and segmentation integrated model trained based on the differentiable diffusion method. Figure 2 This is a cascaded training process. Detailed Implementation

[0044] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0045] Example 1 like Figure 1 As shown, this invention provides an integrated medical image denoising and segmentation model trained based on the differentiable diffusion method, comprising: An improved denoising diffusion module is used to receive the initial medical image, predict the original clean signal of the initial medical image based on the improved denoising diffusion probability model, and obtain the final denoised medical image. The joint motion and segmentation module receives initial medical images, extracts features through a shared encoder, and obtains segmentation data through joint learning via motion estimation and segmentation branches. The segmentation data can be used to quantitatively assess organ function. Cascaded training mechanism: The improved denoising diffusion module and the joint motion and segmentation module are trained end-to-end through a differentiable connection. The joint loss function is used to realize gradient backpropagation and parameter synchronous optimization to complete the cascaded training. The segmentation data obtained by the joint motion and segmentation module after cascaded training is used as the denoised segmented medical image, and the denoised segmented medical image is used as the final output of this method.

[0046] In one embodiment of the present invention, the improved noise reduction and diffusion module includes: The preliminary image denoising N2N unit is used to perform preliminary denoising on the initial medical image based on J-invariance theory to obtain the initial denoised medical image. Traditional preliminary denoising methods often use Gaussian filtering and bilateral filtering, which can only suppress specific types of noise. For mixed and complex noise in medical images, there is a problem of incomplete denoising. This invention combines the Noise2Nose model with J-invariance theory to obtain the preliminary image denoising N2N unit and optimizes the loss function. It can retain the key structural details in the medical image while maintaining the denoising quality, and obtain a clean initial denoised medical image. Meanwhile, the loss function expression for the initial image denoising N2N unit is as follows:

[0047]

[0048]

[0049] in, The loss function for the initial image denoising N2N unit is... For initial denoised medical images, For norm, Denoising function Through aggregation -1 volumetric slices are used as the input image. For initial medical images containing noise, For clean medical imaging, For constant terms, and All are linear coefficients. Additive noise, With a mean of 0 and a covariance matrix of The normal distribution This is a unit matrix consistent with the dimensions of medical imaging.

[0050] The noise level estimation unit is used to calibrate the residual noise based on the residual distribution between the initial denoised medical image and the initial medical image. Based on the residual noise, the minimum timestamp of the diffusion process is calculated using the Markov chain state matching method. The expression for calibrating the residual noise is as follows:

[0051]

[0052]

[0053]

[0054] in, For residual noise, For initial medical images containing noise, Clean medical images for preliminary image denoising N2N unit loss function estimation, for residual noise Perform calibration to ensure the mean value of residual noise. Zero, The adjusted residual noise, For the adjusted denoised medical images, and All are balance coefficients.

[0055] The distance expression for the matching states in a Markov chain is as follows:

[0056] ,

[0057] in, Minimum timestamp For a predefined noise schedule, Input image and predefined noise timetable All possible posterior standard deviations These are hyperparameters and can be adjusted according to the specific application. Let be the probability density function. The initial medical image dimensions are input. It is a natural exponential function. For the residual noise after calibration, To fit the calibrated residual noise to a value with a mean of 0 and a covariance matrix of... Multidimensional Gaussian distribution The distance expression of the matching state of the Markov chain determines the position of the input medical image in the Markov chain, so that subsequent backsampling starts from this matching state of the Markov chain to generate a denoised medical image.

[0058] The U-Net denoising unit with time coding is used to predict the original clean signal of the initial medical image using the minimum timestamp as the time step in a forward propagation manner, resulting in the final denoised medical image. The expression for the final denoised medical image is as follows:

[0059]

[0060]

[0061] in, This is the final denoised medical image obtained after the reverse process of the entire diffusion model. For all possible values ​​of the desired clean, noise-free medical image, This represents the joint probability distribution of medical images during the inverse denoising process. This represents the total time step of the diffusion process. The minimum timestamp is the time step of the current diffusion process. For time step Noisy medical images, The initial noise distribution at the start of the diffusion process. It serves as the core learning component for neural network parameterization of single-step denoising transition probabilities. For mean prediction, For variance prediction, The parameters of the neural network are obtained by training on the medical image dataset; the Markov chain matching state determines the position of the input medical image in the Markov chain, so that subsequent backsampling can start from this state to generate denoised medical images, thereby improving the denoising quality and efficiency.

[0062] To overcome the inherent contradiction between global perception and computational efficiency in current denoising architectures, in one embodiment of the present invention, a denoising unit based on the Mamba architecture is provided to replace the U-Net denoising unit with time coding. In existing technologies, medical image denoising tasks mainly rely on two mainstream architectures: one is U-Net and its variants based on convolutional neural networks (CNNs), and the other is the Transformer architecture based on the self-attention mechanism. The U-Net architecture effectively captures multi-scale features of organs through encoder-decoder structures and skip connections, but its core convolutional operation has a local receptive field, which has inherent limitations in modeling long-range, global semantic dependencies in images, resulting in insufficient ability to restore the consistency of large-scale anatomical structures and remove structural artifacts. The other mainstream architecture, Transformer, can overcome the limitations of locality through the self-attention mechanism and achieve global context modeling, but its computational complexity is proportional to the square of the image sequence length. When processing high-resolution medical images, such as cardiac MRI or ultrasound images, the number of image patches is extremely large, which makes Transformer face huge computational overhead, high memory consumption, and slow training convergence speed, which restricts its feasibility and deployment efficiency in clinical practice and high-dimensional data scenarios.

[0063] The Mamba-based denoising unit, through a unique selective state-space mechanism, can dynamically adjust information propagation based on input content, such as noise patterns and anatomical structures, to prioritize the processing of key information and suppress irrelevant noise. Furthermore, when processing sequence data, the Mamba-based denoising unit requires only linear complexity (O(N)) to reduce computational and memory requirements while maintaining powerful global modeling capabilities. The Mamba-based denoising unit includes: Image segmentation and embedding: The initial denoised medical image as input is , For time step Noisy medical images, The height is the number of pixels. The width in pixels. For the number of channels, Representing grayscale medical images, the initial denoised medical image is segmented into... There are 10 patches, each patch being 1000 pieces. , The total number of image patches, The size of each image patch, i.e. × Pixels, using linear projection to map each patch to dimensional vector middle, This is the initial sequence representation. The feature dimension mapped to each image patch refers to a local region of the image, such as a small P×P block. Dividing the image into multiple patches can transform the image into a series of local regions, making it easier to process the image like a sequence.

[0064] Learnable positional encoding: ,in , For the enhanced sequence, As a learnable location encoding matrix, since the image is divided into blocks and becomes a sequence, and the sequence itself has no location information, the spatial location information of the initial denoised medical image is lost. Therefore, location encoding is needed. To restore location information, It is a learnable matrix with shape [formula missing]. That is, each token (corresponding to a location) has one D Positional encoding of dimensions.

[0065] Time step conditional embedding: Using sinusoidal position embedding to minimize timestamps Mapped to D A dimensional vector, then adjusted for dimensionality through a linear layer: Add a timestamp condition to each token: , Minimum timestamp Here, is a sinusoidal position coding function, and Linear(·) is a linear projection layer. For the final conditional sequence, As a time conditional vector, in order to incorporate time step information into the network, the time step information is encoded into a... D dimensional vector This information is then added to each token, ensuring that each token contains time step information. A token typically refers to an element in the sequence, and each patch is obtained after linear projection. D A dimensional vector is essentially a token; therefore, a token represents the feature representation of a patch, and the entire image is thus represented as a vector composed of... N A sequence of tokens.

[0066] Mamba block: For each Mamba block, the input sequence After layer normalization, through the Mamba layer: Then it goes through a feedforward network (FFN): ,in, For the current layer index, For the first The output sequence of the layer, LayerNorm is the layer normalization operation, Mamba is the forward computation of the selective state-space model, and the input sequence is... After layer normalization, the result is fed into the Mamba layer and then concatenated with the input residual to obtain... Then, after passing through a feedforward network (FFN) and undergoing residual connection again, we obtain... , The output sequence is the Mamba layer, and FFN is a feedforward neural network. For the first The final output sequence of the layer.

[0067] Decoding to an image: Finally, the sequence Rearranged as The feature map is then used to map each token back to its original location using a linear projection. The patch is reassembled into an image. ,in, This represents the total number of Mamba block levels. For the first The final output sequence of the layer, i.e., the final output sequence of the network. For the predicted noisy image; finally, the sequence That is, the output of the last Mamba block, rearranged into a two-dimensional structure, i.e. The feature map is used to map each D-dimensional vector back to a P×P×C block using a linear projection. Finally, these blocks are stitched together to form a complete image of size H×W×C, which is a denoised medical image.

[0068] This invention, through the introduction of the Mamba architecture, addresses the problems of existing CNNs and Transformers in medical image denoising tasks. In terms of performance, the Mamba-based denoising unit, with its global perception and dynamic selective attention capabilities, generates high-quality denoised medical images with more consistent anatomical structures and clearer edges. In terms of efficiency, the linear computational complexity of the Mamba-based denoising unit enables it to efficiently process high-resolution images, significantly reducing computational costs and inference time, providing a foundation for real-time clinical applications. At the system level, the Mamba-based denoising unit performs high-quality denoising, and through differentiable gradient paths in cascaded training, it provides more robust and discriminative feature representations for the joint motion and segmentation modules, thereby stably and significantly improving the accuracy and robustness of medical image segmentation tasks.

[0069] In one embodiment of the present invention, the combined motion and segmentation module specifically includes: A shared encoder is used to perform multi-scale feature extraction and manual annotation on initial medical images to obtain pre-organ movement temporal images, manually annotated pre-organ movement temporal images, post-organ movement temporal images, and manually annotated post-organ movement temporal images. The motion unit is used to stitch together manually annotated images of organ activity before and after movement to predict the deformation field of the activity process between the first and second phase images. The segmentation unit is used to segment manually annotated images of organ activity before and after movement using a fully convolutional network, obtaining segmentation results for the images before and after organ activity.

[0070] In one embodiment of the present invention, the cascaded training mechanism specifically includes: Through a differentiable gradient path, the loss gradient of the joint motion and segmentation module is backpropagated to the improved denoising and diffusion module. The parameters of the improved denoising and diffusion module and the joint motion and segmentation module are jointly supervised through the joint loss function. According to the joint loss function, cascade training is completed. The cascaded joint motion and segmentation module is used to directly output segmentation data with reduced noise interference based on the initial medical image to obtain the final denoised medical image.

[0071] The expression for the joint loss function is as follows:

[0072]

[0073]

[0074]

[0075]

[0076]

[0077]

[0078] in, The loss function for the joint motion and segmentation module, The weights of the loss function are used to estimate the motion. For motion estimation loss function, To divide the weights of the loss function, For the segmentation loss function, For the improved loss function of the denoising diffusion module, For the improved noise reduction and diffusion module, For segmentation modules, Images of organ activity prior to movement. These are phase images of organ activity following movement. Manual annotation of pre-movement phase images of organ activity. Manual annotation of pre-movement phase images of organ activity was obtained. The segmentation results are for the temporal images before organ activity and movement. The segmentation results are for the temporal images after organ activity and movement. These are temporal images of organ activity before movement, after denoising processing. These are temporal images of organ activity after denoising processing. for The spatial x-coordinate of a sampled pixel in an image. This represents the spatial ordinate of the pixel in the image. To predict the displacement in the lateral direction of the deformation field Φ during the activity process between the pre-movement phase image and the post-movement phase image of organ activity, To predict the displacement in the longitudinal direction of the deformation field Φ during the motion process between the pre-movement phase image and the post-movement phase image of organ activity, It is the L2 norm. Here is the Huber loss function.

[0079] like Figure 2 As shown, the specific process of the cascaded training mechanism of this invention is as follows: Data input: Receive raw images of the first phase of organ activity before movement. and its manual annotation Original images of the second phase after motion and its manual annotation This is the original image; An improved denoising diffusion module, connected to the input module, is used to denoise the input data; the improved denoising diffusion module uses timestamps representing the degree of image noise obtained by Markov chain state matching. Using the original image as input, a reverse diffusion process is performed, and the denoised image is output; timestamp The input acts as a coordinator for cascaded training, indirectly aiding in the training of the segmentation task and timestamp. By coordinating the accuracy of the denoising process, the quality of the backpropagation gradient is ensured, enabling the improved denoising diffusion module to effectively guide the joint motion and segmentation module. Ultimately, this allows the joint motion and segmentation module, even when receiving the original image, to output less noisy and more accurate medical image segmentation data. Through precise control of the denoising process, timestamps... The improved denoising diffusion module is ensured to generate a meaningful and computable gradient for its loss function, which is backpropagated through a differentiable path. The ultimate goal is to optimize the joint motion and segmentation module so that it can extract features that are useful for restoring a clean image. The denoised final image output by the improved denoising diffusion module provides an ideal learning target for the joint motion and segmentation module. Through a differentiable gradient path, the joint motion and segmentation module learns robust features that are insensitive to noise and beneficial to the segmentation task. Through the differentiable gradient path, the denoised final image serves as an auxiliary supervision signal, working together with the segmentation loss to optimize the feature extraction basis of the entire system from back to front, thereby enabling the entire system to have stronger segmentation capabilities when faced with the original noisy image.

[0080] A joint motion and segmentation module receives the original image in parallel with the improved denoising and diffusion module; this module includes a shared encoder and a dual decoder structure consisting of a motion estimation branch and a segmentation branch for outputting information on the displacement deformation field Φ and segmentation data; Through a differentiable gradient path, the loss gradient originating from the joint motion and segmentation module is backpropagated to the improved differentiable denoising and diffusion module, thereby integrating the improved denoising and diffusion module with the joint motion and segmentation module. The parameters of the differentiable improved denoising and diffusion module and the joint motion and segmentation module are synchronously supervised by a joint loss function, which is a weighted sum of three sub-losses: denoising loss, motion estimation loss, and segmentation loss. The final output, after optimization by the cascaded training mechanism, is the segmented data with reduced noise interference from the branch output of the joint motion and segmentation module, which serves as the final result of medical image denoising and segmentation.

[0081] By establishing a differentiable gradient path, the improved denoising and diffusion module and the segmentation module were trained collaboratively, enabling joint training of denoising and segmentation to improve segmentation quality and efficiency. Through joint supervision of the joint loss function, the joint motion and segmentation module can directly output segmentation data with noise removed from the original medical images, improving the model's segmentation accuracy and robustness under noise interference and ensuring the reliability of the output results.

[0082] The beneficial effects of this invention are as follows: The denoising diffusion module is improved to make it differentiable, and the improved denoising diffusion module is integrated with the joint motion estimation segmentation module to construct a cascaded model that can be trained end-to-end. This achieves synergistic optimization of denoising and segmentation, solving the problems of traditional denoising methods that damage image structural details due to blindly smoothing noise, and the inability to jointly train the denoising diffusion model with the segmentation model due to its lack of differentiability, resulting in low quality of denoised segmented medical images. This invention improves the accuracy and robustness of medical image segmentation, and by backsampling from the calculated timestamp, it improves the denoising speed and effect. The motion estimation branch and segmentation branch ensure the physiological rationality of the segmentation results, making the medical image segmentation conform to the actual physiological state and achieving higher segmentation quality.

Claims

1. A medical image denoising and segmentation integrated model trained based on the differentiable diffusion method, characterized in that, include: An improved denoising diffusion module is used to receive the initial medical image, predict the original clean signal of the initial medical image based on the improved denoising diffusion probability model, and obtain the final denoised medical image. The joint motion and segmentation module receives initial medical images, extracts features through a shared encoder, and obtains segmentation data through joint learning via motion estimation and segmentation branches. Cascaded training mechanism: The improved denoising diffusion module and the joint motion and segmentation module are trained end-to-end through a differentiable connection. The joint loss function is used to realize gradient backpropagation and parameter synchronous optimization to complete the cascaded training. The segmentation data obtained by the joint motion and segmentation module after cascaded training is used as the denoised segmented medical image.

2. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 1, characterized in that, The improved noise reduction and diffusion module includes: The preliminary image denoising N2N unit is used to perform preliminary denoising on the initial medical image based on the J-invariance theory to obtain the initial denoised medical image. The noise level estimation unit is used to calibrate the residual noise based on the residual distribution between the initial denoised medical image and the initial medical image. Based on the residual noise, the minimum timestamp of the diffusion process is calculated using the Markov chain state matching method. The U-Net denoising unit with time coding is used to predict the original clean signal of the initial medical image by using the minimum timestamp as the time step in a forward propagation manner, so as to obtain the final denoised medical image.

3. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 2, characterized in that, The loss function expression for the initial image denoising N2N unit is as follows: in, The loss function for the initial image denoising N2N unit is... For initial denoised medical images, For norm, Denoising function Through aggregation -1 volumetric slices are used as the input image. For initial medical images containing noise, For clean medical imaging, For constant terms, and All are linear coefficients. Additive noise, The mean is 0 and the covariance matrix is The normal distribution An identity matrix consistent with the dimensions of medical imaging.

4. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 2, characterized in that, The expression for the calibration residual noise is as follows: in, For residual noise, For initial medical images containing noise, Clean medical images for initial image denoising N2N unit loss function estimation. The mean of the residual noise is . The adjusted residual noise, For the adjusted denoised medical images, and All are balance coefficients.

5. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 2, characterized in that, The distance expression for the matching states of the Markov chain is as follows: , in, Minimum timestamp For a predefined noise schedule, Input image and predefined noise timetable All possible posterior standard deviations For hyperparameters, Let be the probability density function. The initial medical image dimensions are input. It is a natural exponential function. For the calibrated residual noise, To fit the calibrated residual noise to a value with a mean of 0 and a covariance matrix of... Multidimensional Gaussian distribution .

6. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 2, characterized in that, The final denoised medical image is expressed as follows: in, This is the final denoised medical image obtained after the reverse process of the entire diffusion model. For all possible values ​​of the desired clean, noise-free medical image, This represents the joint probability distribution of medical images during the inverse denoising process. This represents the total time step of the diffusion process. The minimum timestamp is the time step of the current diffusion process. For time steps Noisy medical images, The initial noise distribution at the start of the diffusion process. It serves as the core learning component for neural network parameterization of single-step denoising transition probabilities. For mean prediction, For variance prediction, These are the parameters of the neural network trained on a medical image dataset.

7. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 2, characterized in that, The time-coded U-Net denoising unit in the improved denoising diffusion module can also be replaced with a Mamba-based denoising unit; the Mamba-based denoising unit is used to process the initial denoised medical image and the minimum timestamp to obtain the denoised medical image. The Mamba-based denoising unit includes: Image segmentation and embedding: Input initial denoised medical image The input initial denoised medical image is segmented into There are 10 patches, and each patch is mapped to a linear projection. dimensional vector In; among them, For time steps Noisy medical images, The height is the number of pixels. The width in pixels. For the number of channels, , The total number of image patches, The size of each image patch This is the initial sequence representation. The feature dimension mapped to each image patch; Learnable positional encoding: for dimensional vector Add position encoding The enhanced sequence is obtained. ;in, , , The learnable positional encoding matrix has the following shape: ; Time step conditional embedding: using sinusoidal position embedding to find the minimum timestamp Mapped to D From the dimensional vector, we obtain the time condition vector. And add the time condition vector to the enhanced sequence. The conditional sequence of the initial denoised medical image is obtained. ;in, , , Minimum timestamp Here, is a sinusoidal position coding function, and Linear(·) is a linear projection layer. This is the conditional sequence for the initial denoised medical images. It is a time condition vector; Mamba block: Conditional sequence for receiving initial denoised medical images. The sequence is then processed sequentially through layer normalization, a Mamba layer, and a feedforward network (FFN) to obtain the final sequence. For each Mamba block, the input sequence is... After layer normalization, through the Mamba layer: Then it goes through a feedforward network FFN: ,in, For the current layer index, For the first The output sequence of the layer, LayerNorm is the layer normalization operation, Mamba is the forward computation of the selective state-space model, and the input sequence is... After layer normalization, the result is fed into the Mamba layer and then concatenated with the input residual to obtain... After passing through a feedforward network FFN and then performing residual connections again, we obtain... , The output sequence is the Mamba layer, and FFN is a feedforward neural network. For the first The final output sequence of the layer, For the final sequence, This represents the total number of Mamba layers in the Mamba block. Decoding to an image: converting the final sequence Rearranged as The feature map is used to map each token back to its feature map via a linear projection. The patch is reconstructed to obtain a denoised medical image. ,in, This is the predicted noisy image.

8. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 1, characterized in that, The joint motion and segmentation module specifically includes: A shared encoder is used to perform multi-scale feature extraction and manual annotation on initial medical images to obtain pre-organ movement temporal images, manually annotated pre-organ movement temporal images, post-organ movement temporal images, and manually annotated post-organ movement temporal images. The motion unit is used to stitch together manually annotated images of organ activity before and after movement to predict the deformation field of the activity process between the first and second phase images. The segmentation unit is used to segment manually annotated images of organ activity before and after movement using a fully convolutional network, obtaining segmentation results for the images before and after organ activity.

9. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 1, characterized in that, The expression for the joint loss function is as follows: in, The loss function for the joint motion and segmentation module, The weights of the loss function are used to estimate the motion. For motion estimation loss function, To divide the weights of the loss function, For the segmentation loss function, For the improved loss function of the denoising diffusion module, For the improved noise reduction and diffusion module, For segmentation modules, Images of organ activity prior to movement. These are phase images of organ activity following movement. Manual annotation of pre-movement phase images of organ activity. Manual annotation of pre-movement phase images of organ activity was obtained. The segmentation results are for the temporal images before organ activity and movement. The segmentation results are for the temporal images after organ activity and movement. These are temporal images of organ activity before movement, after denoising processing. These are temporal images of organ activity after denoising processing. for The spatial x-coordinate of a sampled pixel in an image. This represents the spatial ordinate of the pixel in the image. To predict the displacement in the lateral direction of the deformation field Φ during the activity process between the pre-movement phase image and the post-movement phase image of organ activity, To predict the displacement in the longitudinal direction of the deformation field Φ during the motion process between the pre-movement phase image and the post-movement phase image of organ activity, It is the L2 norm. Here is the Huber loss function.

10. The integrated medical image denoising and segmentation model trained based on the differentiable diffusion method according to claim 9, characterized in that, The cascaded training mechanism specifically includes: Through a differentiable gradient path, the loss gradient of the joint motion and segmentation module is backpropagated to the improved denoising and diffusion module. The parameters of the improved denoising and diffusion module and the parameters of the joint motion and segmentation module are jointly supervised through the joint loss function. Cascaded training is completed according to the joint loss function. The joint motion and segmentation module after cascaded training is used to directly output segmentation data with reduced noise interference based on the initial medical images, resulting in the final denoised medical images.

Citation Information

Cited By

  • Cross-modal medical image synthesis method, system, equipment and medium

    CN121639839A

  • CT-to-MRI and focus integrated generation network training method and application, system, equipment and medium thereof

    CN122065916A

  • A method for training a generative network integrating CT to MRI and lesions, and its application, system, equipment, and media.

    CN122065916B