Image generation method and device, equipment, readable storage medium and program product
The method enhances medical image reconstruction by decoupling and fusing features from multi-modal images, improving precision through Gaussian noise processing and implicit feature extraction.
Patent Information
- Application Number
- CN202510797338.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing medical image reconstruction technology is time-consuming and labor-intensive, with poor stability and versatility, and the accuracy of cross-modal reconstruction images is low, making it difficult to effectively integrate biological information between multimodals, resulting in difficulty in decoupling and expression of anatomical features, and the boundary blurring of the reconstruction images in key areas, affecting the accuracy of quantitative analysis.
By constructing a multimodal feature decoupling network for feature decoupling, using structural prior guidance for feature fusion, and combining implicit diffusion models for denoising, high-quality target mode images are generated to improve reconstruction accuracy.
The feature decoupling and fusion in the hidden space dimension is achieved, which eliminates semantic redundancy, improves the accuracy of image reconstruction and the topological consistency of anatomical structures, and ensures the fidelity of biological structures.
Smart Images

Figure CN120318362A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technologies, and particularly to an image generation method, apparatus, device, readable storage medium, and program product. Background Art
[0002] Medical image reconstruction is an important field in medical image processing. It generates images from raw data through various technologies and algorithms to meet the needs of clinical diagnosis and treatment. The medical image reconstruction technologies in traditional technologies include methods based on atlases, sparse coding, and traditional machine learning. However, these methods are usually time-consuming, laborious, and have poor stability and generality.
[0003] In view of these problems, with the continuous development of deep learning technologies, significant breakthroughs have been made in the field of medical image reconstruction by deep learning. In related technologies, new technologies such as autoencoders (AE), convolutional neural networks (CNN), and generative adversarial networks (GAN) are used to achieve cross-modal image reconstruction. However, the cross-modal reconstruction in related technologies limits the ability to restore image details, resulting in low accuracy of the reconstructed images. Summary of the Invention
[0004] Based on this, it is necessary to provide an image generation method, apparatus, computer device, computer readable storage medium, and computer program product that can improve the accuracy of reconstructed images for the above technical problems.
[0005] In a first aspect, the present application provides an image generation method, including: Obtaining a target modal image that meets a first quality requirement and a reference modal image that meets a second quality requirement; wherein, the first quality requirement is less than the second quality requirement; Performing feature decoupling on the target modal image and the reference modal image to obtain the content feature of the reference modal image and the attribute feature of the target modal image; Fusing the attribute feature and the content feature to obtain a cross-modal fusion feature; Obtaining Gaussian noise, using the cross-modal fusion feature as a conditional input, and performing denoising processing on the Gaussian noise to obtain an implicit feature; Decoding the implicit feature to obtain a second target modal image that meets the second quality requirement.
[0006] In one embodiment, the feature decoupling of the target modality image and the reference modality image to obtain the content feature of the reference modality image and the attribute feature of the target modality image includes: Input the target modality image and the reference modality image into a trained multi-modal feature decoupling network, and perform feature decoupling through the multi-modal feature decoupling network to output the content feature of the reference modality image and the attribute feature of the target modality image.
[0007] In one embodiment, the training method of the multi-modal feature decoupling network includes: Construct a first total loss function of the auto-encoder in the multi-modal feature decoupling network and a second total loss function of the decoupling module in the multi-modal feature decoupling network; Obtain a first sample image set for training the multi-modal feature decoupling network, where the first sample image set includes multiple groups of image groups with at least two quality requirements in different modalities; Register the first image in the first sample image set to the individual space of each first image using a preset image template to obtain respective corresponding multi-channel initial tissue probability maps; Train the encoder according to each first image and the initial tissue probability map corresponding to each first image. When the first function value of the first total loss function is less than a first threshold, complete the training of the encoder; Obtain a second sample image set, and train the decoupling module according to the second sample image set. When the second function value of the second total loss function of the decoupling module is less than a second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
[0008] In one embodiment, the first total loss function includes an image reconstruction loss function and a tissue probability divergence loss function. The training of the auto-encoder according to each first image and the initial tissue probability map corresponding to each first image, and when the first function value of the first total loss function is less than a first threshold, completing the training of the auto-encoder includes: Input each first image into the encoder, output a feature representation, and determine a reconstructed image according to the feature representation input into the decoder; Determine the reconstruction loss value of the image reconstruction loss function according to the reconstructed image and the first image; Determine the target probability distribution and the predicted probability distribution belonging to multiple initial tissue probability maps at the current voxel point in each first image; Determine the tissue probability map divergence loss value of the tissue probability divergence loss function according to the target probability distribution and the predicted probability distribution; Determine a first function value of the first total loss function according to the reconstruction loss value and the tissue probability map divergence loss value. When the first function value is less than a first threshold, complete the training of the encoder.
[0009] In one embodiment, when training the decoupling module according to the second sample image set, when a second function value of a second total loss function of the decoupling module is less than a second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network, including: For each second image in the second sample image set, determine an initial tissue probability map and imaging features corresponding to each second image; Input each second image and its corresponding initial tissue probability map into the trained encoder to obtain initial content features and initial attribute features; Input the initial content features and the initial attribute features into the decoupling module, and perform feature extraction processing on the initial attribute features through an attribute feature extraction network in the decoupling module to obtain target attribute features irrelevant to the content features; Perform enhancement processing on the initial content features through a content feature extraction network in the decoupling module to obtain target content features; Perform feature fusion on the target attribute features and the target content features to obtain new cross-modal fusion features; Determine a second function value of the second total loss function of the decoupling module according to the imaging features, the target attribute features, the target content features, the initial tissue probability map, and the cross-modal fusion features. When the second function value is less than the second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
[0010] In one embodiment, the performing feature extraction processing on the initial attribute features through an attribute feature extraction network in the decoupling module to obtain target attribute features irrelevant to the content features includes: Extract additional content features in the initial attribute features through an attribute feature extraction network in the decoupling module, and delete the additional content features from the initial attribute features to obtain target attribute features irrelevant to the content features; The performing enhancement processing on the initial content features through a content feature extraction network in the decoupling module to obtain target content features includes: Align the initial attribute features and the initial content features, and fuse the additional content features to perform enhancement processing on the initial content features to obtain target content features.
[0011] In one embodiment, each of the second images and their respective initial tissue probability maps are input into a trained encoder to obtain initial content features and initial attribute features, including: Input each of the second images into the encoder to obtain image features, and input the initial tissue probability map corresponding to the second image into the encoder to obtain initial content features; Based on a preset relationship satisfied among the image, the tissue probability map, and the modal attribute map, determine the initial attribute features according to the preset relationship, the image features, and the initial content features.
[0012] In one embodiment, obtaining Gaussian noise, taking the cross-modal fusion feature as a conditional input, and performing denoising processing on the Gaussian noise to obtain implicit features, including: Obtain Gaussian noise, take the Gaussian noise as an input and the cross-modal fusion feature as a conditional input, and input them into a trained implicit diffusion model for progressive denoising to obtain implicit features.
[0013] In one embodiment, the training of the implicit diffusion model includes: Construct a third total loss function for the implicit diffusion model; Obtain a third sample image set for training the implicit diffusion model, where the third sample image set includes multiple groups of image groups with at least one quality requirement; each image group includes at least two modalities; Perform Gaussian blur processing on the third images in the third sample image set to obtain high-frequency detail maps; For each image group in the third sample image set, use a trained multi-modal feature decoupling network to decouple the features of the image group to obtain content features and attribute features, and decode the content features obtained from the reference modal images that meet the second quality requirement to obtain an optimized tissue probability map; Perform progressive noise addition processing on the image features obtained by the third images passing through the encoder to obtain target Gaussian noise; Taking the target Gaussian noise as the starting point for reverse denoising and the cross-modal fusion feature as a control condition, input them into the denoising network of the implicit diffusion model for progressive denoising to obtain the noise prediction for each step of denoising and the finally output predicted image; Determine the function value of the third total loss function according to the noise prediction, the high-frequency detail map, the tissue probability map, the content features, the attribute features, and the predicted image. When the function value of the third total loss function is less than a third threshold, complete the training of the implicit diffusion model to obtain a trained implicit diffusion model.
[0014] In one embodiment, the third image is a second target modality sample image that meets the second quality requirement. The step of performing Gaussian blur processing on the sample image set to obtain a high-frequency detail map includes: Performing Gaussian blur processing on the second target modality sample image to obtain a low-frequency structure image; Processing the target modality high-quality image according to the low-frequency structure image to obtain a high-frequency detail map; Each image group includes a first target modality sample image that meets the first quality requirement and a second reference modality sample image that meets the second quality requirement. The step of using the trained multi-modal feature decoupling network to perform feature decoupling on the image group to obtain content features and attribute features includes: Using the trained multi-modal feature decoupling network to extract content features from the second reference modality sample image, and decoding the content features to obtain an optimized tissue probability map; Using the trained multi-modal feature decoupling network to extract attribute features from the first target modality sample image.
[0015] In one embodiment, the third total loss function includes a noise prediction loss function, a boundary awareness loss function, and a high-frequency texture loss function. The step of determining the function value of the third total loss function according to the noise prediction, the high-frequency detail map, the tissue probability map, the content features, the attribute features, and the predicted image includes: Determining the noise prediction loss value of the noise prediction loss function according to the noise prediction; Calculating the entropy value of each voxel in the predicted image according to the tissue probability map, and determining the error value of each voxel between the predicted image and the second target modality sample image. Determining the boundary awareness loss value of the boundary awareness loss function according to the entropy value and the error value; Determining the high-frequency texture loss value of the high-frequency texture loss function according to the predicted image and the high-frequency detail map; Weighting the noise prediction loss value, the boundary awareness loss value, and the high-frequency texture loss value to determine the function value of the third total loss function.
[0016] In a second aspect, the present application further provides an image generation device, including: A data acquisition module, configured to acquire a target modality image that meets the first quality requirement and a reference modality image that meets the second quality requirement; wherein, the first quality requirement is less than the second quality requirement; A decoupling module for decoupling the features of the target modality image and the reference modality image to obtain the content features of the reference modality image and the attribute features of the target modality image; A feature processing module for fusing the attribute features and the content features to obtain cross-modal fusion features; An image processing module for obtaining Gaussian noise, taking the cross-modal fusion features as conditional inputs, and denoising the Gaussian noise to obtain implicit features; Decoding the implicit features to obtain a second target modality image meeting the second quality requirement.
[0017] In a third aspect, the present application also provides a computer device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above image generation methods are implemented.
[0018] In a fourth aspect, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above image generation methods are implemented.
[0019] In a fifth aspect, the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of any one of the above image generation methods are implemented.
[0020] For the above image generation methods, devices, computer devices, computer-readable storage media, and computer program products, through the target modality image meeting the first quality requirement and the reference modality image meeting the second quality requirement, feature decoupling of the images is realized in the latent space dimension to obtain content features that can be used for cross-modal sharing and attribute features of modal features, eliminating semantic redundancy in the feature space. Through structural prior guidance for feature structures, fusing the decoupled attribute features and content features to obtain cross-modal fusion features, taking the cross-modal fusion features as conditional inputs, denoising Gaussian noise to obtain implicit features; decoding the implicit features to generate a second target modality image meeting the second quality requirement, improving the accuracy of image reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments of the present application or related technologies. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other related drawings can be obtained without creative efforts based on these drawings.
[0022] Figure 1It is an application environment diagram of an image generation method in an embodiment; Figure 2 It is a schematic flowchart of an image generation method in an embodiment; Figure 3 It is a schematic structural diagram of a multi-modal feature decoupling network guided by structural prior in an embodiment; Figure 4 It is a schematic flowchart of a training method for a multi-modal feature decoupling network in an embodiment; Figure 5 It is a schematic diagram of the architecture of a quality-enhanced implicit diffusion model for cross-modal feature fusion in an embodiment; Figure 6 It is a schematic flowchart of a training method for an implicit diffusion model in an embodiment; Figure 7 It is a schematic flowchart of an image generation method in another embodiment; Figure 8 It is a structural block diagram of an image generation device in an embodiment; Figure 9 It is an internal structural diagram of a computer device in an embodiment. Detailed implementation manners
[0023] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0024] The combination of multi-modal brain imaging technology and intelligent imaging analysis methods provides a new objective means for the accurate diagnosis and prognosis evaluation of brain diseases. As the mainstream neuroimaging modalities, magnetic resonance imaging (MRI) and positron emission tomography (PET) have realized the cross-scale pathological feature analysis from macroscopic structure to molecular function through three-dimensional anatomical morphological characterization, glucose metabolism tracing and neurotransmitter dynamic monitoring, and have become the core basis for the clinical diagnosis of neurodegenerative diseases such as Alzheimer's disease and Parkinson's disease.
[0025] However, conventional clinical imaging devices have significant technical bottlenecks: The isotropic spatial resolution of a standard 1.5T MRI system of approximately 1 mm³ makes it difficult to accurately detect microscopic pathological signs such as the volume reduction of the substantia nigra pars compacta, and it has insufficient sensitivity to sub-millimeter pathological changes such as iron ion deposition. Although the ultra-high field 7T MRI can improve the resolution to the order of 0.4 mm³, its clinical application is limited by the scarcity of equipment, the complexity of motion artifact control technology, and the high scanning cost. Although PET imaging can reflect molecular pathological processes such as β-amyloid protein deposition, its spatial resolution is generally lower than 2 mm due to the constraints of gamma photon detection efficiency and tracer pharmacokinetic characteristics, resulting in limited recognition of fine structures.
[0026] The image quality enhancement technology based on deep learning provides a new paradigm for solving the above problems. By establishing a non-linear mapping relationship between the degradation model and high-quality images, this technology can break through the diffraction limit of the physical imaging system and achieve enhanced visualization of sub-voxel pathological features. This not only significantly improves the accuracy and repeatability of quantitative analysis, but also has important clinical value in reducing the risk of radioactive exposure and optimizing the allocation of medical resources, laying a technical foundation for promoting the early screening and precise intervention of neurological diseases.
[0027] In recent years, deep learning has made significant breakthroughs in the field of medical image reconstruction, especially showing important research value in the collaborative processing of multi-modal data. Although the early single-modal mapping model can improve the basic image quality, it fails to effectively integrate the complementary biological information between different modalities, limiting the ability to restore details. Current multi-modal fusion methods mostly adopt the shallow feature splicing strategy, resulting in too high a cross-modal information coupling degree and making it difficult to achieve decoupled expression of anatomical features.
[0028] In the application level of generative models, although the generative adversarial network (GAN) can generate high-fidelity images, its inherent mode collapse problem leads to limited output diversity, and there are defects in the stability of network training. Although the variational autoencoder (VAE) has the advantage of fast sampling, the generated images are often accompanied by blurred edges and texture distortion. Although the diffusion model achieves a balance between image quality and diversity, its iterative generation mechanism leads to low computational efficiency and it is difficult to meet the clinical real-time requirements. It should be noted that medical image reconstruction has strict requirements for the fidelity of anatomical structures, and subtle morphological deviations may cause feature distortion in complex brain regions such as the basal ganglia, which may lead to clinical misjudgment.
[0029] Existing methods generally lack an explicit constraint mechanism for neuroanatomical structures, resulting in blurred boundaries in key regions such as the gray-white matter junction in the reconstructed images, seriously affecting the accuracy of quantitative analysis of brain tissues. This structural distortion problem limits the adaptability of current generation models in the medical field. Therefore, it is necessary to construct a deep representation framework that integrates multi-modal complementary information and introduce anatomical prior constraints to precisely control the reconstruction of brain images, ensure the topological consistency of biological structures, and prevent morphological distortion of key biomarkers (such as iron deposition in the substantia nigra).
[0030] To address the technical problem of low accuracy of the reconstructed images in the above problems, an image generation method is proposed. The image generation method provided in the embodiments of this application can be applied to, for example Figure 1 the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be integrated on the server 104, or placed in the cloud or other network servers. The terminal 102 obtains a target modal image that meets the first quality requirement and a reference modal image that meets the second quality requirement from the server 104; decouples the features of the target modal image and the reference modal image to obtain the content features of the reference modal image and the attribute features of the target modal image; fuses the attribute features and the content features to obtain a cross-modal fusion feature, obtains Gaussian noise, uses the cross-modal fusion feature as a conditional input, and performs denoising processing on the Gaussian noise to obtain implicit features; decodes the implicit features to obtain a second target modal image that meets the second quality requirement.
[0031] Among them, the terminal 102 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0032] In an exemplary embodiment, as Figure 2 shown, an image generation method is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps 202 to step 210. Among them: Step 202, obtain a target modal image that meets the first quality requirement and a reference modal image that meets the second quality requirement; where the index value of the first quality requirement is less than the index value of the second quality requirement.
[0033] Among them, the index value of the quality requirement (i.e., the quality index) can be determined according to a preset resolution or a preset number of pixels. For different quality indices, there is a corresponding preset threshold range, which can be a resolution threshold or a pixel threshold. In this example, the quality index of the image is distinguished by the resolution. The preset threshold range corresponding to the first quality requirement is smaller than the preset threshold range corresponding to the second quality requirement. The quality that meets the first quality requirement can be called low quality, and the quality that meets the second quality requirement can be called high quality.
[0034] The image can be an image in different application scenarios. The image can be a medical image, and the medical image can be a neuroimage. In this example, the neuroimage is taken as an example for illustration. The target modality and the reference modality can refer to the types of images generated by different imaging techniques. The target modality refers to the type of image that is expected to be finally generated or analyzed in multi-modal image analysis. The target modality may be magnetic resonance imaging (MRI) or positron emission tomography (PET), etc. The reference modality refers to other imaging modalities used to assist in the analysis of the target modality. The reference modality may be computed tomography (CT) or ultrasound imaging, etc. If the target modality is MRI, the reference modality can be computed tomography (CT).
[0035] The target modality image and the reference modality image are paired data. The paired data can be understood as image data of different modalities and different quality requirements coming from the same test subject. That is to say, the target modality image and the reference modality image come from the same test subject.
[0036] Exemplarily, obtain a low-quality target modality image I low tar And a high-quality reference modality image I high ref 。
[0037] Step 204: Decouple the features of the target modality image and the reference modality image to obtain the content features of the reference modality image and the attribute features of the target modality image.
[0038] Among them, feature decoupling can be performed using a trained multi-modal feature decoupling network. The multi-modal feature decoupling network is a decoupling network based on structural prior guidance. By constructing a dual-path optimization paradigm of latent space decoupling-recombination, a modal invariance constraint and an anatomical structure fidelity mechanism are established, and the cross-modal shared anatomical structure information (i.e., content features) and modality-specific attribute features (such as texture / contrast, etc.) can be separated from multi-modal neuroimages. Eliminate the distortion of modal attributes and structural content representations caused by the residual anatomical structure information in the attribute encoding; and suppress the unintended carrying of attribute features on the anatomical structure information; thereby improving the biomarker interpretability of the decoupled features in the quantitative analysis of brain diseases.
[0039] The prior guidance of tissue structure can refer to using the known tissue structure information as prior knowledge to guide and optimize the medical image reconstruction process during medical image reconstruction, thereby improving the quality and accuracy of the reconstructed image.
[0040] Exemplarily, the target modality image and the reference modality image are input into the trained multi-modal feature decoupling network, and feature decoupling is performed through the multi-modal feature decoupling network to output the content features of the reference modality image and the attribute features of the target modality image.
[0041] Step 206: Fuse the attribute features and the content features to obtain cross-modal fusion features.
[0042] Among them, feature fusion can be to fuse the decoupled features of multi-modal data through feature modulation of the multi-modal feature decoupling network to obtain cross-modal fusion features.
[0043] Step 208: Obtain Gaussian noise, use the cross-modal fusion features as conditional inputs, and perform denoising processing on the Gaussian noise to obtain implicit features.
[0044] Exemplarily, Gaussian noise is obtained, the Gaussian noise is used as an input, and the cross-modal fusion features are used as conditional inputs and input into the trained implicit diffusion model for step-by-step denoising to obtain implicit features. Among them, the implicit diffusion model uses the cross-modal fusion features as prior conditions, focuses on complex regions with high reconstruction difficulty, and efficiently completes the feature map reconstruction in the low-dimensional implicit space to improve the quality of neuroimages. The complex regions with high reconstruction difficulty can be the junctions of different tissues, such as the gray-white matter junction.
[0045] The implicit diffusion model includes an encoder, a decoder, and a diffusion model. The core of the diffusion model is a conditional neural network, which can be a U-Net structure. The U-Net captures the spatial continuity and local correlation of the image by repeatedly applying convolutional operations.
[0046] It should be noted that considering the different reconstruction difficulties of different regions in neuroimages, the construction of the loss function in the implicit diffusion model in this embodiment is based on a region selection strategy of structure entropy weighted mapping, integrating high-frequency residual analysis and tissue boundary detection methods, constructing a dual anatomical-driven loss function, guiding the model to adaptively focus on complex regions with high reconstruction difficulty, and then focusing on complex regions with high reconstruction difficulty, efficiently completing the feature map reconstruction in the low-dimensional implicit space, and improving the quality of neuroimages.
[0047] Step 210: Decode the implicit features to obtain a second target modality image that meets the second quality requirement.
[0048] Exemplarily, obtain a low-quality target modality image I lowtar and the high-quality reference modality image I high ref 。The obtained low-quality target modality image I low tar and the high-quality reference modality image I high ref are input into the trained multi-modal feature decoupling network for feature decoupling to extract the content features of I high ref and the attribute features of I and I low tar 。Starting from Gaussian noise using the fused cross-modal features as conditions, input into the UNet network based on cross-attention, and gradually denoise to generate implicit features 。The decoder will convert it into a high-quality target modality image I pred 。
[0049] In the above image generation method, for the target modality image meeting the first quality requirement and the reference modality image meeting the second quality requirement, feature decoupling of the images is achieved in the latent space dimension, obtaining content features that can be used for cross-modal sharing and attribute features representing modality features, eliminating semantic redundancy in the feature space. Feature decoupling is guided by structural priors, and based on the fused attribute features and content features obtained by decoupling, cross-modal fusion features are obtained. Using the cross-modal fusion features as conditions for input, Gaussian noise is denoised to obtain implicit features; the implicit features are decoded to generate a second target modality image meeting the second quality requirement, improving the accuracy of image reconstruction.
[0050] Considering that in the related art, the brain imaging feature decoupling method based on image space decomposition has significant limitations: its content encoder has suboptimal capture ability for brain anatomical structure features, and some anatomical information remains in the attribute feature space, resulting in semantic mixing of cross-domain (content domain and attribute domain) representations of neuroimages. Therefore, a multi-modal feature decoupling network based on structural prior guidance is used to separate content and attributes of the input multi-modal data in the feature latent space.
[0051] In an exemplary embodiment, feature decoupling of the target modality image and the reference modality image to obtain the content features of the reference modality image and the attribute features of the target modality image includes: inputting the target modality image and the reference modality image into the trained multi-modal feature decoupling network, performing feature decoupling through the multi-modal feature decoupling network, and outputting the content features of the reference modality image and the attribute features of the target modality image.
[0052] Such as Figure 3As shown in the figure, it is a schematic diagram of the structure of a multi-modal feature decoupling network guided by structural prior. The multi-modal feature decoupling network includes an encoder, a decoder, a self-attention mechanism module, a cross-attention mechanism module, and feature modulation. Among them, the input is image data of different modalities with different quality requirements, as well as the tissue probability map corresponding to the image data. Among them, the image input to the model is the image data, and the other three input images are the initial tissue probability maps. For example, they can respectively represent the initial tissue probability maps of three tissues: cerebrospinal fluid, gray matter, and white matter. Based on Figure 3 the network architecture diagram shown, a training method for the multi-modal feature decoupling network is provided, as Figure 4 shown, including steps 402 to 410, where: Step 402, construct the first total loss function of the autoencoder in the multi-modal feature decoupling network, and the second total loss function of the decoupling module in the multi-modal feature decoupling network.
[0053] It can be understood that the multi-modal feature decoupling network includes two major parts: an autoencoder and a decoupling module. Among them, the autoencoder includes an encoder and a decoder. The encoder encodes the image into features, and the decoder reconstructs the image using the features. During training here, not only the encoder needs to be trained, but also the decoder needs to be trained.
[0054] Among them, the first total loss function can also be called the reconstruction loss function, which is used to train the autoencoder in the multi-modal feature decoupling network. The first total loss function includes an image reconstruction loss function and a tissue probability divergence loss function. The image reconstruction loss function is used to minimize the difference between the original input and the reconstructed output. The image reconstruction loss function can be expressed as: (1) Among them, x includes images with different quality requirements under different modalities, and may include the first target modality sample image I low tar that meets the first quality requirement and the first reference modality sample image I low ref , the second target modality sample image I high tar that meets the second quality requirement and the second reference modality sample image I high ref .
[0055] The reconstruction loss L recon represents the sum of the reconstruction errors of all sample images, and ∑ x sums over all samples x . Enc( x ) takes the input data xConvert to the latent space representation. The decoder function Dec(Enc( x )) converts the latent space representation back to the original data space, attempting to reconstruct the input data x . The L2 norm is used to calculate the distance between two vectors. Here, it calculates the square root of the sum of the squared differences between the original input x and the reconstructed data Dec(Enc( x ))
[0056] The tissue probability divergence loss function is used to compare the similarity between the probability distribution predicted by the model and the true distribution, thereby optimizing the training of the model. Specifically, it is used to compare the target probability distribution and the predicted probability distribution of the current voxel point on the sample image belonging to the corresponding multi-channel tissue probability map. In this example, considering that the tissue probability map represents the probability distribution of N tissues, the Kullback-Leibler (KL) divergence is used to calculate the loss. The tissue probability divergence loss function L kl can be expressed as: (2) where q label and q pred represent the target probability distribution and the predicted probability distribution of the current voxel point v on the sample image belonging to N tissues, respectively
[0057] According to the image reconstruction loss function and the tissue probability divergence loss function, the first total loss function L ae can be expressed as: (3) where is the first preset coefficient
[0058] The decoupling module can be a two-stream depth feature decoupling module, including an attribute feature optimization branch (i.e., an attribute feature branch network) and a content feature optimization branch (i.e., a content feature branch network). With this two-stream depth feature decoupling architecture, the initially obtained content and attribute features are further optimized and reconstructed, effectively generating a tissue probability map with complete anatomical information, while ensuring that the attribute feature space strictly strips content-related interference
[0059] The second total loss function of the decoupling module includes an implicit self-reconstruction loss function, an implicit cross-reconstruction loss function, a content decoupling loss function, an attribute decoupling loss function, a contrast decoupling loss function, and an attribute consistency loss function
[0060] Among them, the implicit self-reconstruction loss function is used to characterize the difference between the updated attribute features and content features obtained by the decoupling module from the image features and image features of the sample image, and the two features will generate a new difference between the image features after passing through the feature modulation module (Modula). Implicit self-reconstruction loss function can be expressed as: (4) y includes the high- and low-quality image features of the target modality and the reference modality, L L1 norm is to calculate the sum of the absolute values of the vector elements, represents the sample image features of the model y and the difference between the value calculated by the feature modulation Modula. Here, Modula may be a specific function or operation, and are the attribute features and content features output by the decoupling module from the image features of the sample image.
[0061] The implicit cross-reconstruction loss function is used to characterize the loss between the high- and low-quality image features of different modalities and the new image features obtained after the exchange and combination of the updated attribute features and content features obtained by the decoupling module. Implicit cross-reconstruction loss function can be expressed as Equation (5), where 、 respectively represent the low-quality image features of the target modality calculated by the feature modulation Modula and the updated attribute features obtained by the decoupling module, 、 、 respectively represent the high-quality image features of the target modality calculated by the feature modulation Modula, the updated attribute features obtained by the decoupling module, and the updated content features obtained by the decoupling module, 、 respectively represent the low-quality image features of the reference modality calculated by the feature modulation Modula and the updated attribute features obtained by the decoupling module, 、 、 respectively represent the high-quality image features of the reference modality calculated by the feature modulation Modula, the updated attribute features obtained by the decoupling module, and the updated content features obtained by the decoupling module: (5) Content decoupling loss function: Constraining the similarity between the content features of different modalities and different qualities and the initial probability map features, as shown in Equation (6), where represents the initial probability map features, and respectively represent the updated content features obtained by the decoupling module from the low-quality target modality image and the updated content features obtained by the decoupling module from the low-quality reference modality image, represents the Charbonnier loss.
[0062] (6) Attribute decoupling loss function: Forcing the same-modal attribute features to be consistent and maximizing the differences in different-modal attribute features, as shown in Equation (7), where and respectively represent the updated attribute features obtained by the decoupling module from the low-quality target modality image and the updated attribute features obtained by the decoupling module from the high-quality target modality image, and respectively represent the updated attribute features obtained by the decoupling module from the low-quality reference modality image and the updated attribute features obtained by the decoupling module from the high-quality reference modality image.
[0063] (7) Contrast decoupling loss function: Enhancing the feature decoupling ability through contrast learning, as shown in Equation (8).
[0064] (8) Attribute consistency loss function: Constraining the local smoothness of the attribute feature map, making the gradient change of the attribute feature map as small as possible during the model optimization process, so as to achieve the effect of local smoothness, which helps to generate more natural and reasonable neuroimages. In Equation (9), ▽ represents the horizontal gradient and the vertical gradient, λ is a coefficient for balancing the smoothness degree, and y includes the high- and low-quality image features of the target modality and the reference modality.
[0065] (9) After the autoencoder is trained, fix the network parameters of the autoencoder, and then train the two-stream deep feature decoupling module. The second total loss function of this module is as shown in Equation (10): (10) Step 404, obtain the first sample image set for training the multi-modal feature decoupling network, where the first sample image set includes multiple groups of image groups with at least two quality requirements under different modalities.
[0066] Among them, different modalities can include the target modality and the reference modality, and the quality requirements can include the first quality requirement and the second quality requirement. That is to say, there are two target modality images that meet the first quality requirement and the second quality requirement under the target modality, and there are two reference modality images that meet the first quality requirement and the second quality requirement under the reference modality.
[0067] It should be noted that the training of the neural network is based on the obtained training sample data and is trained in batches (batch) to calculate the loss function. That is to say, taking the first sample image set as the training sample set, the training sample set can be divided into multiple small batches, and the neural network model can be iteratively trained based on each small batch. Then, a small batch includes several first sample image sets, and thus also includes multiple first target modality sample images, first reference modality sample images, second target modality sample images, and second reference modality sample images.
[0068] Step 406: Use a preset image template to register the first image in the first sample image set to the individual space of each first image, and obtain the corresponding multi-channel initial tissue probability map for each.
[0069] Among them, the first image includes the first target modality sample image that meets the first quality requirement, the second target modality sample image that meets the second quality requirement, the first reference modality sample image that meets the first quality requirement, and the second reference modality sample image that meets the second quality requirement.
[0070] For each training, it can be to register the sample image used for training to the individual space of each sample image using a preset image template, and obtain the corresponding multi-channel initial tissue probability map for each.
[0071] For example, use a standard brain template (such as MNI152) to register to the target modality individual space to generate a 3-channel tissue probability map P, which respectively corresponds to the probability distributions of cerebrospinal fluid (CSF), gray matter (GM), and white matter (WM). Here, the tissue probability map P is the initial tissue probability map. That is to say, at this time, only the initial content features of the sample image can be obtained, and the initial attribute features cannot be directly determined.
[0072] Step 408: Train the autoencoder according to each first image and the initial tissue probability map corresponding to the first image. When the first function value of the first total loss function is less than the first threshold, the training of the autoencoder is completed.
[0073] Exemplarily, for each training, use the sample image used for training and the initial tissue probability map corresponding to the sample image to train the autoencoder, calculate the first total loss function value of the autoencoder, adjust the parameters of the autoencoder according to the first total loss function value, and then perform the next batch of training. When the first function value of the first total loss function of the autoencoder is less than the first threshold, the training of the autoencoder is completed.
[0074] Further, in an exemplary embodiment, the autoencoder is trained according to each first image and the initial tissue probability map corresponding to each first image. When the first function value of the first total loss function is less than the first threshold, the training of the autoencoder is completed, including: Input each first image into the encoder to output a feature representation, and determine a reconstructed image according to the feature representation input into the decoder; determine the reconstruction loss value of the image reconstruction loss function according to the reconstructed image and the first image; determine the target probability distribution and the predicted probability distribution belonging to multiple initial tissue probability maps at the current voxel point in each first image; determine the tissue probability map divergence loss value of the tissue probability divergence loss function according to the target probability distribution and the predicted probability distribution; determine the first function value of the first total loss function according to the reconstruction loss value and the tissue probability map divergence loss value. When the first function value is less than the first threshold, the training of the autoencoder is completed.
[0075] Exemplarily, based on the specific calculation methods of the image reconstruction loss function and the tissue probability divergence loss function described above, determine the reconstruction loss value of the image reconstruction loss function according to the reconstructed image and the first image; determine the target probability distribution and the predicted probability distribution belonging to multiple initial tissue probability maps at the current voxel point in each first image; determine the tissue probability map divergence loss value of the tissue probability divergence loss function according to the target probability distribution and the predicted probability distribution. Determine the first function value of the first total loss function according to the reconstruction loss value and the tissue probability map divergence loss value.
[0076] Step 410, obtain a second sample image set, and train the decoupling module according to the second sample image set. When the second function value of the second total loss function of the decoupling module is less than the second threshold, the training of the decoupling module is completed, and a trained multi-modal feature decoupling network is obtained.
[0077] Among them, when the training of the autoencoder is completed, the network parameters of the autoencoder are fixed. Obtain a second sample image set from the prepared training sample data to train the decoupling module, that is, train the two-stream depth feature decoupling module. The second sample image set can be exactly the same as or partially the same as the first sample image set, and no specific limitation is made here.
[0078] Exemplarily, train the decoupling module according to the initial content features and initial attribute features determined by the trained encoder and the sample images. When the function value of the second total loss function of the decoupling module is less than the second threshold, the training of the decoupling module is completed, and a trained multi-modal feature decoupling network is obtained.
[0079] In an exemplary embodiment, training the decoupling module according to the second sample image set, when the second function value of the second total loss function of the decoupling module is less than the second threshold, the training of the decoupling module is completed, and a trained multi-modal feature decoupling network is obtained, including: For each second image in the second sample image set, determine the initial tissue probability map and imaging features corresponding to each second image; input each second image and its corresponding initial tissue probability map into the trained encoder to obtain the initial content features and initial attribute features; input the initial content features and initial attribute features into the decoupling module, and perform feature extraction processing on the initial attribute features through the attribute feature extraction network in the decoupling module to obtain target attribute features that are irrelevant to the content features; perform enhancement processing on the initial content features through the content feature extraction network in the decoupling module to obtain target content features; perform feature fusion on the target attribute features and target content features to obtain new cross-modal fusion features; determine the second function value of the second total loss function of the decoupling module according to the imaging features, target attribute features, target content features, initial tissue probability map, and cross-modal fusion features. When the second function value is less than the second threshold, complete the training of the decoupling module to obtain the trained multi-modal feature decoupling network.
[0080] It can be understood that based on the above method for determining the initial tissue probability map, only the initial content features of the sample images can be obtained, while the initial attribute features cannot be directly determined.
[0081] In an exemplary embodiment, inputting each second image and its corresponding initial tissue probability map into the trained encoder to obtain the initial content features and initial attribute features includes: inputting each second image into the encoder to obtain imaging features, inputting the initial tissue probability map corresponding to the second image into the encoder to obtain the initial content features; based on the preset relationship satisfied among the image, tissue probability map, and modal attribute map, determine the initial attribute features according to the preset relationship, imaging features, and initial content features.
[0082] It can be understood that the sample image I can be decomposed into a tissue probability map P and a modal attribute map A, that is, the preset relationship satisfied among the image, tissue probability map, and modal attribute map is I = A⊙P, which can be specifically expressed as: (11) On this basis, input the sample image and its corresponding initial tissue probability map into their respective encoders, output the imaging features and initial content features. On this basis, perform a transformation based on the relationship satisfied by formula (11) in the image space to obtain the relationship satisfied by the attribute features, content features, and imaging features in the feature space, and then the initial attribute features can be determined. The initial attribute feature f A 、the initial content feature f P and the imaging feature f I The relationship satisfied in the feature space can be expressed as: (12) Among them, τ is an extremely small constant to avoid a zero denominator.
[0083] Among them, determining the function value of the second total loss function of the decoupling module according to the image feature, target attribute feature, target content feature, and initial probability map includes: Based on the calculation methods of the above implicit self-reconstruction function, implicit cross-reconstruction loss function, content decoupling loss function, attribute decoupling loss function value, contrast decoupling loss function, and attribute consistency loss function, according to the cross-modal fusion feature and the image feature of the target modal sample image in the second sample image set, determine the implicit self-reconstruction function value of the implicit self-reconstruction function; according to the image feature determined by the second sample image set and the cross-modal fusion feature, determine the implicit cross-reconstruction loss value of the implicit cross-reconstruction loss function; according to the target content features with different modality and different quality requirements and their corresponding initial probability map features, determine the content decoupling loss value of the content decoupling loss function; according to the difference value between the initial attribute features with different quality requirements in the same modality, determine the attribute decoupling loss value of the attribute decoupling loss function value; according to the target content features and target attribute features with different modality and different quality requirements, obtain the contrast decoupling loss value of the contrast decoupling loss function; according to the target content features and target attribute features with different modality and different quality requirements, perform attribute consistency loss calculation to obtain the attribute consistency loss value of the attribute consistency loss function. Determine the function value of the second total loss function of the decoupling module based on these loss function values.
[0084] In the above method, based on the multi-modal feature decoupling network guided by structural prior, a two-stream deep feature decoupling architecture is constructed. Through the multi-resolution cross-modal joint optimization strategy, the decoupled representation of neuroimages is realized in the latent space dimension, and the content and attributes are separated in the feature latent space. And based on the feature modulation method, an anatomical structure-guided cross-modal feature fusion paradigm is established, that is, a dual-path optimization paradigm of latent space decoupling-recombination. By extracting rich anatomical structure prior knowledge through a pre-trained high-quality auxiliary modality encoder, extracting modal attribute features from the target modal image, using the feature modulation method to fuse the decoupled features of multi-modal data, and through the cross-attention mechanism, realizing precise conditional control to ensure the structure-attribute consistency of the generated image, which helps to reduce the distortion phenomenon.
[0085] Considering that the obtained initial attribute features are determined based on the registration method, the obtained initial attribute features are not accurate enough. Therefore, other features in the initial attribute features need to be deleted.
[0086] Optionally, in an exemplary embodiment, the initial attribute features are processed by the attribute feature extraction network in the decoupling module to obtain target attribute features that are irrelevant to the content features, including: extracting the additional content features in the initial attribute features through the attribute feature extraction network in the decoupling module, and deleting the additional content features from the initial attribute features to obtain target attribute features that are irrelevant to the content features.
[0087] Exemplarily, based on Figure 3 the shown network architecture diagram, on the attribute feature optimization branch, a self-attention module (SA) is used to further extract the content information in the initial modal attribute features, and the additional content features are deleted from the original attribute features to generate content-irrelevant attribute features, that is, target attribute features.
[0088] Correspondingly, the initial content features are enhanced by the content feature extraction network in the decoupling module to obtain target content features, including: aligning the initial attribute features and the initial content features, and fusing the additional content features to enhance the initial content features to obtain target content features.
[0089] Exemplarily, cross-attention is used to align the two features, and the additional content information extracted on the attribute feature optimization branch is fused into the content feature optimization branch to generate enhanced content features, that is, target content features, by supplementing details.
[0090] Furthermore, after the processing of the two branches, the decoupling network (module) can generate target content features rich in anatomical structure-related details , and target attribute features that only represent the modality and are irrelevant to the content . and are input into the input feature modulation module for feature fusion to obtain new image features, which are then input into the decoder to reconstruct the neuroimage.
[0091] In this way, compared with directly decoupling the image in the prior art, decoupling the image features and using the two-branch optimization method to further enhance the decoupling on the basis of the initial decoupling can effectively suppress the information leakage between the different feature components obtained by decomposition and improve the interpretability representation ability of the neuroimage features.
[0092] It is understandable that the implicit diffusion model progressively adds noise to the implicit representation during the forward diffusion process and realizes high-quality sample generation based on the reverse denoising process. Compared with the traditional pixel-level generation paradigm, the implicit diffusion model utilizes the compact representation characteristics of the low-dimensional feature space, significantly reducing the computational complexity of high-dimensional data reconstruction. By extracting semantically consistent deep features from noisy multi-modal images as diffusion priors, the anatomical structure fidelity of the generated samples can be effectively improved. More importantly, such models support precise regulation of the generation process through the conditional embedding mechanism, which provides a theoretical framework for the collaborative reconstruction of multi-modal neuroimages.
[0093] In an exemplary embodiment, Gaussian noise is obtained, and the cross-modal fusion feature is used as a conditional input to perform denoising processing on the Gaussian noise to obtain an implicit feature, including: Obtain Gaussian noise, use the Gaussian noise as the input, and the cross-modal fusion feature as the conditional input, and input them into the trained implicit diffusion model for progressive denoising to obtain the implicit feature.
[0094] Among them, the trained implicit diffusion model is a quality-enhanced implicit diffusion model for cross-modal feature fusion. As Figure 5 shown, the schematic diagram of the architecture of the quality-enhanced implicit diffusion model for cross-modal feature fusion includes an encoder, a decoder, forward noise addition, and reverse denoising. Among them, the output of the diffusion process is pure Gaussian white noise feature data, which is also the input of the denoising process. The white cuboid represents the representation of the image in the feature implicit space, that is, the image feature. This model extracts modality-shared content features (for example, representing the spatial distribution of organizational structures such as gray matter, white matter, and cerebrospinal fluid) and modality-specific attribute features (for example, describing texture characteristics such as the contrast of different modality images) through a multi-modal feature decoupling network; then, the cross-modal fusion feature obtained by integrating the decoupled features of different modalities through the feature modulation module is used as a condition and injected into the denoising prediction network of the diffusion model to achieve quality-enhanced reconstruction of neuroimages under the constraint of rich anatomical structure information.
[0095] Based on Figure 5 the schematic diagram of the architecture of the quality-enhanced implicit diffusion model, a training method for the implicit diffusion model is provided. As Figure 6 shown, it includes the following steps: Step 602, construct the third total loss function of the implicit diffusion model.
[0096] It should be noted that in order to enable the model to adaptively focus on complex regions with high reconstruction difficulty, a structure fidelity enhancement strategy is determined. The structure fidelity enhancement strategy includes: designing a region selection strategy based on structure entropy weighted mapping, integrating high-frequency residual analysis and tissue boundary detection methods, and constructing a dual anatomical driving loss function, that is, the third total loss function.
[0097] The third total loss function includes a noise prediction loss function, a boundary awareness loss function, and a high-frequency texture loss function. During the training of the implicit diffusion model, the noise for each denoising step is predicted. The obtained noise prediction can be determined based on the UNet architecture using the cross-attention mechanism to fuse conditional information. The conditional information can be the content features after decoupling the reference modality high-quality image and the attribute features after decoupling the target modality low-quality image determined by decoupling the sample image set using the above-trained multi-modal feature decoupling network. The noise prediction formula can be expressed as: (13) where \(Z_t\) represents the feature data after adding noise at the \(t\)-th time step, is represented as the content features after decoupling the reference modality high-quality image, is represented as the attribute features after decoupling the target modality low-quality image.
[0098] Based on the above noise prediction formula, the noise prediction loss function can be determined as: (14) On this basis, to better evaluate the structural fidelity of the quality-enhanced reconstructed image, we calculate the entropy value of each voxel using the tissue probability map based on the reference modality. The entropy value reflects the uncertainty of the voxel tissue attribution. The higher the entropy value, the more likely it is that the voxel is located in the boundary region of different tissue structures. Therefore, when calculating the boundary awareness loss, the model should pay more attention to high-entropy voxels to improve the reconstruction quality of the boundary region.
[0099] The boundary awareness loss function can be expressed as: (15) The entropy value \(H(x, y, z)\) can be expressed as: (16) where \(N\) is the number of tissue categories, i.e., the number of tissue probability maps, and \(P i is the probability of the \(i\)-th category. represents the predicted image, represents the real sample image.
[0100] It can be understood that calculating the structural entropy using the tissue probability map decoupled from the high-quality data of the auxiliary modality as the weight of the reconstruction loss function helps the model actively focus on the tissue boundary region. On the one hand, the high-quality image of the auxiliary modality is used, and the result is more reliable; on the other hand, each voxel point is assigned a corresponding weight, which solves the alignment problem.
[0101] The high-frequency texture loss function guides the network to better restore the texture information of the image by calculating the difference between the reconstructed image and the real image in the texture details. The high-frequency texture loss function can be expressed as: (17) where I T label represents the texture detail image of the real image, which can also be called the high-frequency detail map, containing the high-frequency information in the image, such as the edges of the structures in the image, small structures, etc.; I T pred represents the texture detail image of the predicted image, that is, the predicted high-frequency detail map.
[0102] On the basis of the completion of the training of the multi-modal feature decoupling network guided by the structural prior, train the quality-enhanced implicit diffusion model for cross-modal feature fusion. The total loss function of this model is shown in Equation (18): (18) Step 604: Obtain a third sample image set for training the implicit diffusion model. Among them, the third sample image set includes multiple groups of image groups with at least one quality requirement; each image group includes at least two modalities.
[0103] Among them, the third sample image set is also determined by the sample image set for accurate training. The third sample image set can be completely the same as, partially the same as, or different from the first image sample set and the second image sample set.
[0104] Step 606: Perform Gaussian blur processing on the third image of the third sample image set to obtain a high-frequency detail map.
[0105] Among them, the third image can be a high-quality target modality image. The high-frequency detail map refers to obtaining the high-frequency component by subtracting the corresponding Gaussian blurred version from the high-quality target modality image. The determination method of the high-frequency detail map includes: performing Gaussian blur processing on the second target modality sample image to obtain a low-frequency structure image; processing the high-quality image of the target modality according to the low-frequency structure image to obtain a high-frequency detail map.
[0106] Exemplarily, by adding Gaussian blur noise to the original image (such as a high-quality target modality image) and performing Gaussian blur on the original image, a low-frequency structure image is obtained. Then, the texture and details can be calculated by subtracting the low-frequency structure image from the original image. This is because Gaussian blur will smooth the image and remove high-frequency information. Then, subtracting the blurred image from the original image mainly leaves the high-frequency texture and detail parts.
[0107] Step 608: For each image group in the third sample image set, use the trained multi-modal feature decoupling network to decouple the features of the image group to obtain content features and attribute features, and decode the content features obtained from the reference modal images that meet the second quality requirement to obtain an optimized tissue probability map.
[0108] Among them, the image group includes sample images of different qualities in different modalities. For example, it can include second reference modal sample images, first target modal sample images, and second target modal sample images. Among them, the second reference modal sample image can also be called a high-quality reference modal image, and the first target modal sample image can also be called a low-quality target modal image.
[0109] Exemplarily, use the trained multi-modal feature decoupling network to extract features from the second reference modal sample images in each image group to obtain content features, and decode the content features to obtain an optimized tissue probability map; use the trained multi-modal feature decoupling network to extract features from the first target modal sample images to obtain attribute features.
[0110] Furthermore, use the trained multi-modal feature decoupling network to extract content features from the reference modal high-quality images in the sample image set, input the content features obtained from the reference modal images that meet the second quality requirement into the decoder to obtain an optimized tissue probability map. And use the trained multi-modal feature decoupling network to extract content features from the reference modal high-quality images and attribute features from the target modal low-quality images. Then input the attribute features of the target modal and the content features of the reference modal into the feature modulation module to obtain cross-modal fusion features, and use them as prior conditions to inject into the denoising prediction network of the diffusion model.
[0111] Step 610: Gradually add noise to the image features obtained from the third image through the encoder to obtain target Gaussian noise.
[0112] Among them, the third image can be a high-quality target modal image, that is, a target modal image that meets the second quality requirement.
[0113] Exemplarily, determine the image features obtained from the target modal images that meet the second quality requirement in the third sample image set, perform forward noise addition on the image features, and gradually add Gaussian noise in the low-dimensional implicit feature space (rather than the original image space) to obtain pure Gaussian white noise feature data, that is, target Gaussian noise. Among them, the time step t ∈ {1,..., T}. The noise schedule uses a cosine scheduler to balance training stability and reconstruction quality.
[0114] Step 612: Using the target Gaussian noise as the starting point for reverse denoising and the cross-modal fusion features as the control condition, input them into the denoising network of the implicit diffusion model for progressive denoising to obtain the noise prediction for each denoising step and the finally output predicted image.
[0115] Exemplarily, based on the above forward noise addition for reverse denoising, the reverse denoising uses the target Gaussian noise as the starting point for reverse denoising and the cross-modal fusion features as the control condition. Based on the UNet architecture, the cross-attention mechanism is used to fuse the conditional information to obtain the noise prediction for each denoising step and the finally output predicted image. For example, the content features (anatomical structure) decoupled from the high-quality reference modal image and the attribute features (modal characteristics) decoupled from the low-quality target modal image are fused to obtain the fused high-quality features as the condition to achieve precise control of the implicit diffusion model for brain image quality enhancement.
[0116] Step 614: Determine the function value of the third total loss function according to the noise prediction, high-frequency detail map, tissue probability map, content features, attribute features, and predicted image. When the function value of the third total loss function is less than the third threshold, the training of the implicit diffusion model is completed, and the trained implicit diffusion model is obtained.
[0117] It can be understood that during the training process of the above diffusion model, based on the constructed loss function above, calculate the total loss function value during the training process, and update the model parameters of the diffusion model according to the total loss function.
[0118] Exemplarily, respectively based on the above noise prediction loss function, boundary perception loss function, and high-frequency texture loss function, determine the noise prediction loss value of the noise prediction loss function according to the noise prediction; calculate the entropy value of each voxel in the predicted image according to the tissue probability map, and determine the error value of each voxel between the predicted image and the second target modal sample image, and determine the boundary perception loss value of the boundary perception loss function according to the entropy value and the error value; determine the high-frequency texture loss value of the high-frequency texture loss function according to the predicted image and the high-frequency detail map; weight the noise prediction loss value, boundary perception loss value, and high-frequency texture loss value to determine the function value of the third total loss function. When the function value of the third total loss function is less than the third threshold, the training of the implicit diffusion model is completed, and the trained implicit diffusion model is obtained.
[0119] Among them, when determining the high-frequency texture loss value of the high-frequency texture loss function according to the predicted image and the high-frequency detail map, it is necessary to determine the predicted high-frequency detail map of the predicted image through the determination method of the high-frequency detail map described above. On this basis, based on the above formula (17), determine the high-frequency texture loss value of the high-frequency texture loss function according to the predicted high-frequency detail map and the high-frequency detail map.
[0120] It should be noted that for the above quality-enhanced implicit diffusion model for cross-modal feature fusion, in view of medical images and the problem of blurred reconstruction in micro-structurally complex regions such as the basal ganglia and hippocampus in existing methods, a dual anatomy-driven loss function is proposed: on the one hand, a boundary-aware loss based on structural entropy weighting is constructed using the tissue probability map of the reference modality, where the reconstruction difficulty of key boundaries such as gray matter-cerebrospinal fluid is quantified by calculating the information entropy; on the other hand, a high-frequency texture consistency loss is designed, a high-frequency detail map is constructed using Gaussian blur, and the reconstruction accuracy of complex textures in sub-structural regions such as thalamic nuclei is enhanced by comparing the target with the generated result.
[0121] The above quality-enhanced implicit diffusion model for cross-modal feature fusion realizes the efficient fusion of cross-modal features in the low-dimensional latent space by constructing a content-attribute decoupled guided conditional generation mechanism; uses a loss function system with anatomical interpretability, that is, uses high-frequency detail maps and tissue probability maps as constraints for the model loss function, and through a weight allocation mechanism, guides the model to adaptively focus on complex regions with high reconstruction difficulty, so that the reconstructed images meet the diagnostic-level quality standards in terms of tissue boundaries and texture fidelity.
[0122] It should be noted that the coefficient λ of the first total loss function, the second total loss function, and the third total loss function all take values in the range of 0-1.
[0123] In an exemplary embodiment, as Figure 7 shown, an image generation method is provided. Taking the case where this method is applied to the Figure 1 terminal as an example, it includes the following steps 702 to step 716. Among them: Step 702, obtain a first sample image set and a second sample image set for training, and train a multi-modal feature decoupling network according to the first sample image set and the second sample image set to obtain a trained multi-modal feature decoupling network.
[0124] Step 704, obtain a third sample image set, and determine the high-frequency detail map of the second target modality image and the tissue probability map of the second reference modality sample image in the third sample image set.
[0125] Step 706, use the trained multi-modal feature decoupling network to extract features from the third sample image set to obtain the attribute features of the first target modality image and the content features of the second reference modality image.
[0126] Exemplarily, using the trained multi-modal feature decoupling network described above, content features and attribute features are respectively extracted from the high-quality reference modal images and the low-quality target modal images; then, the attribute features of the target modality and the content features of the reference modality are input into the feature modulation module to obtain cross-modal fusion features, which are used as prior conditions and injected into the denoising prediction network of the diffusion model.
[0127] Step 708: Generate control conditions for the implicit diffusion model based on the attribute features and content features.
[0128] Step 710: Add noise to the third sample image set in the forward direction to obtain target noise, and perform reverse denoising based on the control conditions to train the implicit diffusion model, obtaining a trained implicit diffusion model.
[0129] Step 712: Obtain the image to be processed, which includes the target modal image to be processed that meets the first quality requirement and the reference modal image to be processed that meets the second quality requirement.
[0130] Step 714: Perform feature decoupling on the sample image through the trained multi-modal feature decoupling network to obtain the attribute features of the target modal image to be processed and the content features of the reference modal image to be processed.
[0131] Step 716: Starting from Gaussian noise, using the fused cross-modal features as conditions, input them into the cross-attention-based UNet network of the implicit diffusion model, gradually denoise to generate implicit features, and decode the implicit features to obtain high-quality target modal images.
[0132] Among them, the fused cross-modal features are determined by fusing the attribute features and content features. It should be noted that the specific implementation method in this embodiment can be implemented in the manner defined above, and will not be elaborated here.
[0133] In this embodiment, a multi-modal feature decoupling-generation collaborative architecture driven by structural prior is constructed. Aiming at the core problems in the multi-modal neuroimage quality enhancement task, such as cross-modal feature redundancy, high computational complexity of high-dimensional reconstruction, and anatomical structure distortion, a two-stream deep feature decoupling architecture is constructed to separate neuroimage features into cross-modal shared anatomical structure information (content features) and modality-specific attribute features, effectively eliminating semantic redundancy in the feature space; and a cross-modal conditional generation framework is designed. By designing a conditional diffusion model in the feature latent space and using a content-attribute two-stream guidance mechanism and a feature modulation method, cross-modal feature fusion is achieved. On this basis, a region selection strategy based on structural entropy weighted mapping is designed, integrating high-frequency residual analysis and tissue boundary detection methods, and a dual anatomical-driven loss function is constructed to guide the model to adaptively focus on complex regions with high reconstruction difficulty. That is to say, in this embodiment, the quality of the image is enhanced through three dimensions: a multi-modal feature decoupling mechanism, a cross-modal conditional generation framework, and a structural fidelity enhancement strategy, avoiding the problem of image structural distortion and ensuring the topological consistency of biological structures.
[0134] It should be understood that although the steps in the flowcharts involved in the above embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0135] Based on the same inventive concept, an embodiment of the present application also provides an image generation device for implementing the above-mentioned image generation method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the image generation device provided below can refer to the limitations on the image generation method in the above text, and will not be repeated here.
[0136] In an exemplary embodiment, as Figure 8 shown, an image generation device is provided, including: a data acquisition module 802, a decoupling module 804, a feature processing module 806, and an image processing module 808, where: The data acquisition module 802 is configured to acquire a target modality image that meets the first quality requirement and a reference modality image that meets the second quality requirement; wherein, the first quality requirement is less than the second quality requirement.
[0137] The decoupling module 804 is configured to perform feature decoupling on the target modality image and the reference modality image to obtain the content features of the reference modality image and the attribute features of the target modality image.
[0138] The feature processing module 806 is configured to fuse the attribute features and the content features to obtain cross-modal fusion features.
[0139] The image processing module 808 is configured to obtain Gaussian noise, use the cross-modal fusion features as conditional input to perform denoising processing on the Gaussian noise to obtain implicit features; decode the implicit features to obtain a second target modality image that meets the second quality requirement.
[0140] In the above image generation device, by performing feature decoupling of the target modality image that meets the first quality requirement and the reference modality image that meets the second quality requirement in the latent space dimension, the content features that can be used for cross-modal sharing and the attribute features of the modality features are obtained, semantic redundancy in the feature space is eliminated, the feature structure is guided by the structural prior, and the cross-modal fusion features are obtained by fusing the decoupled attribute features and content features. Using the cross-modal fusion features as conditional input, Gaussian noise is denoised to obtain implicit features; the implicit features are decoded to generate a second target modality image that meets the second quality requirement, improving the accuracy of image reconstruction.
[0141] In an exemplary embodiment, the decoupling module 804 is further configured to input the target modality image and the reference modality image into a trained multi-modal feature decoupling network, perform feature decoupling through the multi-modal feature decoupling network, and output the content features of the reference modality image and the attribute features of the target modality image.
[0142] In an exemplary embodiment, the above image generation device further includes a training module, and the training module is configured to construct a first total loss function of the autoencoder in the multi-modal feature decoupling network and a second total loss function of the decoupling module in the multi-modal feature decoupling network; Obtain a first sample image set for training the multi-modal feature decoupling network, where the first sample image set includes multiple groups of image groups with at least two quality requirements in different modalities; For the first image in the first sample image set, use a preset image template to register it to the individual space of each first image to obtain respective corresponding multi-channel initial tissue probability maps; Train the encoder according to each first image and the initial tissue probability map corresponding to each first image. When the first function value of the first total loss function is less than the first threshold, the training of the autoencoder is completed; Obtain a second sample image set, train a decoupling module according to the second sample image set. When the second function value of the second total loss function of the decoupling module is less than a second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
[0143] In an exemplary embodiment, the training module is further configured to input each first image into an encoder, output a feature representation, and determine a reconstructed image according to the feature representation through a decoder; Determine a reconstruction loss value of an image reconstruction loss function according to the reconstructed image and the first image; Determine a target probability distribution and a predicted probability distribution belonging to multiple initial tissue probability maps at a current voxel point in each first image; Determine a tissue probability map divergence loss value of a tissue probability divergence loss function according to the target probability distribution and the predicted probability distribution; Determine a first function value of a first total loss function according to the reconstruction loss value and the tissue probability map divergence loss value. When the first function value is less than a first threshold, complete the training of the autoencoder.
[0144] In an exemplary embodiment, the training module is further configured to, for each second image in the second sample image set, determine an initial tissue probability map and an image feature corresponding to each second image; Input each second image and its corresponding initial tissue probability map into the trained encoder to obtain an initial content feature and an initial attribute feature; Input the initial content feature and the initial attribute feature into the decoupling module, and perform feature extraction processing on the initial attribute feature through an attribute feature extraction network in the decoupling module to obtain a target attribute feature independent of the content feature; Perform enhancement processing on the initial content feature through a content feature extraction network in the decoupling module to obtain a target content feature; Perform feature fusion on the target attribute feature and the target content feature to obtain a new cross-modal fusion feature; Determine a second function value of the second total loss function of the decoupling module according to the image feature, the target attribute feature, the target content feature, the initial tissue probability map, and the cross-modal fusion feature. When the second function value is less than a second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
[0145] In an exemplary embodiment, the training module is further configured to extract additional content features in the initial attribute feature through an attribute feature extraction network in the decoupling module, and delete the additional content features from the initial attribute feature to obtain a target attribute feature independent of the content feature; Performing enhancement processing on the initial content feature through a content feature extraction network in the decoupling module to obtain a target content feature, including: Align the initial attribute features and the initial content features, and fuse the additional content features to enhance the initial content features, obtaining the target content features.
[0146] In an exemplary embodiment, the training module is further configured to input each second image into an encoder to obtain image features, and input the initial tissue probability map corresponding to the second image into the encoder to obtain the initial content features; Based on the preset relationship satisfied among the image, the tissue probability map, and the modal attribute map, determine the initial attribute features according to the preset relationship, the image features, and the initial content features.
[0147] In an exemplary embodiment, the image processing module 808 is further configured to obtain Gaussian noise, use the Gaussian noise as the input and the cross-modal fusion features as the conditional input, and input them into the trained implicit diffusion model for progressive denoising to obtain implicit features.
[0148] In an exemplary embodiment, the training module is further configured to construct a third total loss function for the implicit diffusion model; Obtain a third sample image set for training the implicit diffusion model, where the third sample image set includes multiple groups of image groups with at least one quality requirement; each image group includes at least two modalities; Perform Gaussian blur processing on the third images in the third sample image set to obtain high-frequency detail maps; For each image group in the third sample image set, use the trained multi-modal feature decoupling network to decouple the features of the image group to obtain content features and attribute features, and decode the content features obtained from the reference modal images that meet the second quality requirement to obtain an optimized tissue probability map; Perform progressive noise addition processing on the image features obtained by the third images passing through the encoder to obtain target Gaussian noise; Use the target Gaussian noise as the starting point for reverse denoising and the cross-modal fusion features as the control condition, and input them into the denoising network of the implicit diffusion model for progressive denoising to obtain the noise prediction for each step of denoising and the finally output predicted image; Determine the function value of the third total loss function according to the noise prediction, the high-frequency detail map, the tissue probability map, the content features, the attribute features, and the predicted image. When the function value of the third total loss function is less than the third threshold, complete the training of the implicit diffusion model to obtain the trained implicit diffusion model.
[0149] In an exemplary embodiment, the training module is further configured to perform Gaussian blur processing on the second target modal sample images to obtain low-frequency structure images; Process the target modal high-quality images according to the low-frequency structure images to obtain high-frequency detail maps; Each image group includes a first target modality sample image that meets the first quality requirement and a second reference modality sample image that meets the second quality requirement. The trained multi-modal feature decoupling network is used to decouple the features of the image group to obtain content features and attribute features, including: The trained multi-modal feature decoupling network is used to extract features from the second reference modality sample image to obtain content features, and decode the content features to obtain an optimized tissue probability map; The trained multi-modal feature decoupling network is used to extract features from the first target modality sample image to obtain attribute features.
[0150] In an exemplary embodiment, the training module is further configured to determine the noise prediction loss value of the noise prediction loss function according to the noise prediction; Calculate the entropy value of each voxel in the predicted image according to the tissue probability map, and determine the error value of each voxel between the predicted image and the second target modality sample image. Determine the boundary awareness loss value of the boundary awareness loss function according to the entropy value and the error value; Determine the high-frequency texture loss value of the high-frequency texture loss function according to the predicted image and the high-frequency detail map; Weight the noise prediction loss value, the boundary awareness loss value, and the high-frequency texture loss value to determine the function value of the third total loss function.
[0151] Each module in the above image generation device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0152] In an exemplary embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 9As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, a mobile cellular network, near field communication (NFC), or other technologies. The computer program, when executed by the processor, implements an image generation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the computer device housing, or an external keyboard, touchpad, or mouse, etc.
[0153] Those skilled in the art can understand that Figure 9 the structure shown in the figure is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0154] In one embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0155] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0156] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0157] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0158] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, artificial intelligence (AI) processors, etc., and are not limited thereto.
[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this application.
[0160] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of this application. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. An image generation method, characterized in that, The method includes: Obtaining a target modal image meeting the first quality requirement and a reference modal image meeting the second quality requirement; wherein, the first quality requirement is less than the second quality requirement; Performing feature decoupling on the target modal image and the reference modal image to obtain the content feature of the reference modal image and the attribute feature of the target modal image; Fusing the attribute feature and the content feature to obtain a cross-modal fusion feature; Obtaining Gaussian noise, using the cross-modal fusion feature as a conditional input, and performing denoising processing on the Gaussian noise to obtain an implicit feature; Decoding the implicit feature to obtain a second target modal image meeting the second quality requirement.
2. The method according to claim 1, characterized in that, The performing feature decoupling on the target modal image and the reference modal image to obtain the content feature of the reference modal image and the attribute feature of the target modal image includes: Inputting the target modal image and the reference modal image into a trained multi-modal feature decoupling network, performing feature decoupling through the multi-modal feature decoupling network, and outputting the content feature of the reference modal image and the attribute feature of the target modal image.
3. The method according to claim 2, wherein The training method of the multi-modal feature decoupling network includes: Constructing a first total loss function of the auto-encoder in the multi-modal feature decoupling network and a second total loss function of the decoupling module in the multi-modal feature decoupling network; Obtaining a first sample image set for training the multi-modal feature decoupling network, wherein the first sample image set includes image groups with at least two quality requirements in different modalities; Registering the first image in the first sample image set to the individual space of each first image using a preset image template to obtain respective corresponding multi-channel initial tissue probability maps; Training the auto-encoder according to each first image and the initial tissue probability map corresponding to each first image, and when the first function value of the first total loss function is less than a first threshold, completing the training of the auto-encoder; Obtaining a second sample image set, training the decoupling module according to the second sample image set, and when the second function value of the second total loss function of the decoupling module is less than a second threshold, completing the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
4. The method according to claim 3, characterized in that The first total loss function includes an image reconstruction loss function and a tissue probability divergence loss function. The training the auto-encoder according to each first image and the initial tissue probability map corresponding to each first image, and when the first function value of the first total loss function is less than a first threshold, completing the training of the auto-encoder includes: Inputting each first image into an encoder, outputting a feature representation, and determining a reconstructed image according to the feature representation input into a decoder; Determining a reconstruction loss value of the image reconstruction loss function according to the reconstructed image and the first image; Determining a target probability distribution and a predicted probability distribution belonging to multiple initial tissue probability maps at the current voxel point in each first image; Determine the tissue probability map divergence loss value of the tissue probability divergence loss function according to the target probability distribution and the predicted probability distribution; Determine the first function value of the first total loss function according to the reconstruction loss value and the tissue probability map divergence loss value. When the first function value is less than the first threshold, complete the training of the autoencoder.
5. The method according to claim 4, wherein Training the decoupling module according to the second sample image set. When the second function value of the second total loss function of the decoupling module is less than the second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network, including: For each second image in the second sample image set, determine the initial tissue probability map and image features corresponding to each second image; Input each second image and its corresponding initial tissue probability map into the trained encoder to obtain initial content features and initial attribute features; Input the initial content features and the initial attribute features into the decoupling module, and perform feature extraction processing on the initial attribute features through the attribute feature extraction network in the decoupling module to obtain target attribute features that are independent of the content features; Perform enhancement processing on the initial content features through the content feature extraction network in the decoupling module to obtain target content features; Perform feature fusion on the target attribute features and the target content features to obtain new cross-modal fusion features; Determine the second function value of the second total loss function of the decoupling module according to the image features, the target attribute features, the target content features, the initial tissue probability map, and the cross-modal fusion features. When the second function value is less than the second threshold, complete the training of the decoupling module to obtain a trained multi-modal feature decoupling network.
6. The method according to claim 5, characterized in that, The performing feature extraction processing on the initial attribute features through the attribute feature extraction network in the decoupling module to obtain target attribute features that are independent of the content features includes: Extract the additional content features in the initial attribute features through the attribute feature extraction network in the decoupling module, and delete the additional content features from the initial attribute features to obtain target attribute features that are independent of the content features; The performing enhancement processing on the initial content features through the content feature extraction network in the decoupling module to obtain target content features includes: Align the initial attribute features and the initial content features, and fuse the additional content features to perform enhancement processing on the initial content features to obtain target content features.
7. The method according to claim 5, wherein The inputting each second image and its corresponding initial tissue probability map into the trained encoder to obtain initial content features and initial attribute features includes: Input each second image into the encoder to obtain image features, and input the initial tissue probability map corresponding to the second image into the encoder to obtain initial content features; Based on the preset relationship satisfied among the image, the tissue probability map, and the modal attribute map, determine the initial attribute features according to the preset relationship, the image features, and the initial content features.
8. The method according to any one of claims 3 to 7, characterized in that Obtaining Gaussian noise, taking the cross-modal fusion feature as a conditional input, and performing denoising processing on the Gaussian noise to obtain an implicit feature, including: Obtaining Gaussian noise, taking the Gaussian noise as an input and the cross-modal fusion feature as a conditional input, and inputting them into a trained implicit diffusion model for step-by-step denoising to obtain an implicit feature.
9. The method according to claim 8, wherein The training of the implicit diffusion model includes: Constructing a third total loss function of the implicit diffusion model; Obtaining a third sample image set for training the implicit diffusion model, where the third sample image set includes multiple groups of image groups with at least one quality requirement; each image group includes at least two modalities; Performing Gaussian blur processing on the third images in the third sample image set to obtain high-frequency detail maps; For each image group in the third sample image set, using a trained multi-modal feature decoupling network to decouple the features of the image group to obtain content features and attribute features, and decoding the content features obtained from the reference modal images that meet the second quality requirement to obtain an optimized tissue probability map; Performing step-by-step noise addition processing on the image features obtained by the third images passing through the encoder to obtain target Gaussian noise; Taking the target Gaussian noise as the starting point for reverse denoising and the cross-modal fusion feature as a control condition, and inputting them into the denoising network of the implicit diffusion model for step-by-step denoising to obtain the noise prediction for each step of denoising and the finally output predicted image; Determining the function value of the third total loss function according to the noise prediction, the high-frequency detail map, the tissue probability map, the content features, the attribute features, and the predicted image. When the function value of the third total loss function is less than the third threshold, the training of the implicit diffusion model is completed to obtain a trained implicit diffusion model.
10. The method according to claim 9, characterized in that, The third image is a second target modal sample image that meets the second quality requirement. The performing Gaussian blur processing on the sample image set to obtain high-frequency detail maps includes: Performing Gaussian blur processing on the second target modal sample image to obtain a low-frequency structure image; Processing the high-quality target modal image according to the low-frequency structure image to obtain a high-frequency detail map; Each image group includes a first target modal sample image that meets the first quality requirement and a second reference modal sample image that meets the second quality requirement. The using a trained multi-modal feature decoupling network to decouple the features of the image group to obtain content features and attribute features includes: Using a trained multi-modal feature decoupling network to extract features from the second reference modal sample image to obtain content features, and decoding the content features to obtain an optimized tissue probability map; Using the trained multi-modal feature decoupling network to extract features from the first target modal sample image to obtain attribute features.
11. The method according to claim 10, wherein The third total loss function includes a noise prediction loss function, a boundary awareness loss function, and a high-frequency texture loss function. Determining the function value of the third total loss function according to the noise prediction, the high-frequency detail map, the tissue probability map, the content feature, the attribute feature, and the predicted image includes: Determining a noise prediction loss value of the noise prediction loss function according to the noise prediction; Calculating an entropy value of each voxel in the predicted image according to the tissue probability map, and determining an error value of each voxel between the predicted image and the second target modality sample image, and determining a boundary awareness loss value of the boundary awareness loss function according to the entropy value and the error value; Determining a high-frequency texture loss value of the high-frequency texture loss function according to the predicted image and the high-frequency detail map; Weighting the noise prediction loss value, the boundary awareness loss value, and the high-frequency texture loss value to determine the function value of the third total loss function.
12. An image generation device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire a target modality image meeting a first quality requirement and a reference modality image meeting a second quality requirement; wherein, the first quality requirement is less than the second quality requirement; A decoupling module, configured to perform feature decoupling on the target modality image and the reference modality image to obtain the content feature of the reference modality image and the attribute feature of the target modality image; A feature processing module, configured to fuse the attribute feature and the content feature to obtain a cross-modal fusion feature; An image processing module, configured to acquire Gaussian noise, use the cross-modal fusion feature as a conditional input, and perform denoising processing on the Gaussian noise to obtain an implicit feature; Decoding the implicit feature to obtain a second target modality image meeting the second quality requirement.
13. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Magnetic resonance brain image super-resolution reconstruction method and device, equipment and storage medium
CN117649344A
Image generation method and device
CN117934653A
Image segmentation method and device, computer equipment and storage medium
CN117974693A
Medical image generation method and device, electronic equipment and storage medium
CN119579708A
Cited By
Low-dose PET image denoising method based on prompt learning and line integral projection constraint
CN120953115A
Low dose pet image denoising method based on prompt learning and line integral projection constraint
CN120953115B