Cross-modal image style migration method and system and medium

Through a cross-modal style transfer method with dual encoder structure and multiple combined loss functions, combined with feature space diffusion model, the problem of pixels and anatomical structure misalignment in cross-modal style transfer in the medical image field is solved, and high-quality image conversion and generation effects are achieved.

CN120147108AActive Publication Date: 2025-06-13HANGLOK-TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202411623387.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-06-13
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

The prior art has problems with pixel and anatomical structure misalignment in the cross-modal style transfer in the field of medical images, resulting in low image conversion accuracy, unclear and inaccurate image conversion.

Method used

A cross-modal style transfer method with a dual encoder structure and multiple combined loss functions is adopted to achieve the alignment of cross-modal medical images at the feature level through characterization learning, and a feature space diffusion model is used to dynamically adjust the feature space during noise addition and denoising to improve the generation performance and efficiency.

Benefits of technology

It effectively reduces the problems caused by pixel and anatomical structure misalignment, realizes high-quality cross-modal medical image style transfer, and improves the clarity and accuracy of generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147108A_ABST
    Figure CN120147108A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-modal image style migration method and system and a medium, and is used for converting an image of a first modal type into an image of a second modal type, and the method comprises the steps: obtaining an initial image of the first modal type of an object, and inputting the initial image into a pre-constructed cross-modal style migration model; the establishment of the model comprises the following steps: respectively inputting first and second modal type images in a sample into a first encoder and a second encoder to correspondingly obtain first and second image features; training the basic model by using a preset total loss function, so that the trained model can align the first image feature and the second image feature; the total loss function is at least related to similarity loss, and the similarity loss is determined according to the first image feature and the second image feature; and the trained model extracts the initial image by using the first encoder to obtain a target image feature, and reconstructs a target image of a second modal type based on the target image feature. According to the invention, the precision, definition and accuracy of cross-modal image conversion can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of medical devices and computer vision technologies, and in particular, to a cross-modal image style transfer method, system, and medium. Background Art

[0002] In interventional therapy, for liver space-occupying lesions lacking typical liver cancer imaging features, it is expected to obtain a clear pathological diagnosis through liver lesion puncture biopsy ultrasound technology or under CT guidance. Although the update of surgical imaging equipment and technology has improved the accuracy and efficiency during the puncture process, due to the relatively low signal-to-noise ratio of intraoperative CT imaging, the uncertainty during the puncture process has increased.

[0003] Computed tomography (CT) imaging technology is a non-invasive imaging method that can provide detailed cross-sectional images and is widely used in clinical observation, disease diagnosis, and treatment guidance. Enhanced-phase CT injects iodine contrast agent into the patient's body to improve the visibility of specific tissues or blood vessels, thereby significantly enhancing the contrast and clarity of the images. This technology helps doctors more accurately identify lesions or other abnormalities [1] However, due to limitations in the use of contrast agents and other factors during the surgical process, non-contrast enhanced CT scans are usually used as the main imaging method. Therefore, image generation models based on convolutional neural networks have great potential for development in realizing cross-modal medical image style transfer.

[0004] The emergence of generative adversarial networks (GANs) and diffusion models has laid a strong theoretical foundation for image generation and opened up new possibilities for cross-modal style transfer of medical images. In recent years, the medical applications based on these generative models have increased significantly, such as CycleGAN [2] 、Pix2pix [3] 、DDPM [4] 、DDIM [5] and other technologies have received extensive attention and applications. In addition, in order to achieve the lightweight of the model, the combination of GAN and diffusion models has gradually come into the view of researchers, such as LDM [6] ,ALDM [7] These generative models can not only capture and reconstruct complex image features, but also optimize the quality of the generated images for style transfer through complex model design and theoretical derivation.

[0005] Currently, the mainstream image generation models at home and abroad usually use Stable Diffusion [6]As the basic framework. However, there are certain limitations in the application of Stable Diffusion in the field of medical images. First of all, in medical image processing, images are often three-dimensional (such as CT or MRI scans). However, Stable Diffusion essentially processes two-dimensional image data, and challenges will be encountered when directly applied to three-dimensional images. Secondly, the generation of medical images not only requires the clarity of the images, but also the accurate reconstruction of anatomical structures and lesion characteristics. However, it is difficult for Stable Diffusion to meet these requirements under the condition of unconditional guidance.

[0006] The relevant literatures [1] to [7] in the prior art cited above are as follows:

[0007] [1] Hansen N J. Computed tomographic angiography of the abdominal aorta[J]. Radiologic Clinics, 2016, 54(1): 35-54.

[0008] [2] Zhu J Y, Park T, Isola P, et al. Unpaired image-to-image translation using cycle-consistent adversarial networks[C] / / Proceedings of the IEEE international conference on computer vision. 2017: 2223-2232.

[0009] [3] Isola P, Zhu J Y, Zhou T, et al. Image-to-image translation with conditional adversarial networks[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1125-1134.

[0010] [4] Ho J, Jain A, Abbeel P. Denoising diffusion probabilistic models[J]. Advances in neural information processing systems, 2020, 33: 6840-6851.

[0011] [5] Song J, Meng C, Ermon S. Denoising diffusion implicit models[J]. arXiv preprint arXiv:2010.02502, 2020.

[0012] [6] Rombach R, Blattmann A, Lorenz D, et al. High-resolution image synthesis with latent diffusion models[C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2022:10684-10695.

[0013] [7] Kim J, Park H. Adaptive latent diffusion model for 3d medical image to image translation: Multi-modal magnetic resonance imaging study[C] / / Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision. 2024:7604-7613.

[0014] The disclosure of the above background art is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of this patent application, nor will it necessarily provide technical guidance. Without clear evidence indicating that the above content was publicly available before the filing date of this patent application, the above background art should not be used to evaluate the novelty and inventiveness of this application. Summary of the Invention

[0015] The object of the present invention is to provide a cross-modal image style transfer method, system and medium, which can effectively reduce problems such as low conversion accuracy, blurriness and inaccuracy of cross-modal images caused by misalignment of pixels and anatomical structures in cross-modal data.

[0016] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0017] A cross-modal image style transfer method for converting an image of a first modality type into an image of a second modality type, the method comprising the following steps:

[0018] Obtain an initial image of the first modality type of an object, and input the initial image into a pre-constructed cross-modal style transfer model;

[0019] Wherein, the cross-modal style transfer model is pre-established through the following steps:

[0020] Collect a learning sample set, and each sample in the learning sample set includes an image of the first modality type and an image of the second modality type of the same object;

[0021] Design a basic model, the basic model includes a first encoder and a second encoder; wherein, the first encoder and the second encoder are configured to extract features from the image;

[0022] Use the learning sample set to optimize and train the basic model in the following manner:

[0023] Input the image of the first modality type in the sample into the first encoder to obtain a first image feature; and input the image of the second modality type in the sample into the second encoder to obtain a second image feature;

[0024] Use a preset total loss function to train the basic model, so that the first image feature extracted by the first encoder of the trained model is aligned with the second image feature extracted by the second encoder, to obtain the cross-modal style transfer model; the total loss function is at least related to the similarity loss, and the similarity loss is determined according to the first image feature and the second image feature;

[0025] The cross-modal style transfer model uses the first encoder therein to extract a target image feature from the initial image, and reconstructs a target image of the second modality type based on the target image feature.

[0026] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the basic model further includes a quantization module, and during the optimization training process of the basic model, the quantization module is configured to perform quantization processing on the first image feature to obtain a quantization feature; the total loss function is further at least related to the quantization loss, and the quantization loss is determined according to the first image feature and the quantization feature;

[0027] The cross-modal style transfer model uses the quantization module therein to perform quantization processing on the target image feature to obtain a target quantization feature, and reconstructs a target image of the second modality type based on the target quantization feature.

[0028] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the calculation formula of the quantization loss is expressed as follows:

[0029]

[0030] Among them, L quan represents the quantization loss, and E c represents the second image feature, and detach() means that during the model training process, the gradient backpropagation of the training is paused. represents the quantization feature.

[0031] Furthermore, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the base model further includes a decoder. During the optimization training process of the base model, the decoder is configured to reconstruct an image of the second modality type based on the second image feature to obtain a reconstructed image of the second modality type; the total loss function is further at least related to the reconstruction loss, and the reconstruction loss is based on the second modality type image and the reconstructed image of the second modality type;

[0032] The cross-modal style transfer model uses the decoder therein to reconstruct a target image of the second modality type based on the target image feature.

[0033] Furthermore, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the calculation formula of the reconstruction loss is expressed as follows:

[0034]

[0035] Among them, L rec represents the reconstruction loss, x c represents the second modality type image, represents the reconstructed image of the second modality type.

[0036] Furthermore, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the base model further includes a classifier. During the optimization training process of the base model, the classifier is configured to determine a classification loss, and the classification loss is determined based on the second image and the reconstructed image of the second modality type reconstructed based on the second image feature; the total loss function is further at least related to the classification loss.

[0037] Furthermore, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the calculation formula of the classification loss is expressed as follows:

[0038]

[0039] Among them, L dis represents the classification loss, D(·) represents the classifier, and x c represents the second modality type image, Represents the reconstructed image of the second modal type.

[0040] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the similarity loss is determined according to the first image feature and the second image feature, and the calculation formula of the similarity loss is expressed as follows:

[0041] L sim =||E c -E n || 2

[0042] Among them, L sim represents the similarity loss, E c represents the second image feature, and E n represents the first image feature.

[0043] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the base model further includes a feature optimizer configured to optimize the features extracted from the image by the first encoder;

[0044] The cross-modal style transfer model uses the feature optimizer therein to optimize the target image features to obtain optimized features, and reconstructs a target image of the second modal type based on the optimized features.

[0045] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the feature optimizer includes a 3D convolutional layer, a normalization layer, and an activation function layer connected in sequence.

[0046] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the method further includes converting the target image features into target fusion features by using a pre-constructed feature space diffusion model; the cross-modal style transfer model reconstructs a target image of the second modal type based on the target fusion features;

[0047] Among them, the feature space diffusion model is pre-established through the following steps:

[0048] Collect a diffusion learning sample set, and each sample in the diffusion learning sample set includes the first image feature and the second image feature;

[0049] Design a base diffusion model, which includes a dynamic similarity mask module and a diffusion module, and the dynamic similarity mask module is configured to obtain a mask of its input image features;

[0050] Use the diffusion learning sample set to optimize and train the base diffusion model in the following manner:

[0051] Input the first image feature and the second image feature into the dynamic similarity mask module to obtain a dynamic similarity mask for the first image feature and the second image feature;

[0052] Input the first image feature and the dynamic similarity mask into the diffusion module to obtain a first fused feature;

[0053] And train the basic diffusion model using a preset diffusion loss function, so that the first fused feature output by the trained model incorporates the structural information of the second image feature on the basis of the first image feature, obtaining the feature space diffusion model; the preset diffusion loss function is at least related to the dynamic similarity mask;

[0054] The feature space diffusion model uses the diffusion module therein to incorporate the structural information of the image feature of the second modality type image on the basis of the target image feature to obtain a first target fused feature, and determines the target fused feature according to the first target fused feature.

[0055] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the basic diffusion model further includes a noise processing module, and the noise processing module is configured to perform noise addition processing and / or noise removal processing on the first fused feature to obtain a second fused feature; the diffusion loss function is further at least related to the input noise and output noise of the basic diffusion model, the input noise is determined by the first image feature and the second image feature, and the output noise is determined by the second fused feature;

[0056] The feature space diffusion model uses the noise processing module therein to perform noise addition and / or noise removal processing on the first target fused feature to obtain a second target fused feature; and determines the second target fused feature as the target fused feature.

[0057] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the diffusion loss function is expressed by the following formula:

[0058]

[0059] where DSM represents the dynamic similarity mask, E n represents the first image feature, E c represents the second image feature, represents the dot product operation, represents the random noise ∈ sampled from the standard normal distribution (N(0,1)) at each time step t, combined with the input data E n and Ec Perform multiple samplings, and the output noise ∈ θ The expected error between the input noise, i.e., the true noise ∈, where ∈ represents the input noise, ∈ θ represents the output noise, represents the square of the L2 norm.

[0060] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, the calculation formula of the dynamic similarity mask is as follows:

[0061]

[0062] where DSM represents the dynamic similarity mask, τ represents the current training cycle number, Total epochs represents the total number of training cycles, <·,·> represents the cosine similarity between the first image feature and the second image feature, and the value range of <·,·> is [-1, 1], transform the value range of the similarity to [0, 1], min[·,·] represents taking the minimum value between two elements, and α is a scaling factor, which is a preset value.

[0063] Further, based on any one of the foregoing technical solutions or a combination of multiple technical solutions, it further includes: determining the first target fusion feature as the target fusion feature.

[0064] According to another aspect of the present invention, the present invention provides a cross-modal image style transfer system, and the cross-modal image style transfer system uses the cross-modal image style transfer method as described in any one of the foregoing technical solutions or a combination of multiple technical solutions to convert an image of a first modal type into an image of a second modal type.

[0065] According to another aspect of the present invention, the present invention provides a computer-readable storage medium for storing program instructions, and the program instructions are configured to be called to execute the steps of the method as described in any one of the foregoing technical solutions or a combination of multiple technical solutions.

[0066] The beneficial effects brought by the technical solutions provided by the present invention are as follows:

[0067] a. In the process of converting an image of the first modality type into an image of the second modality type by the cross-modal image style transfer method provided by the present invention, cross-modal medical images are aligned at the feature level through representation learning, thus effectively reducing the problems caused by misalignment of pixels and anatomical structures in cross-modal data in practical applications. Specifically, the cross-modal image style transfer method cleverly decomposes the style transfer task from an image of the first modality type to an image of the second modality type into two subtasks: one is to achieve alignment of the features of the image of the first modality type and the image of the second modality type; the other is to complete the reconstruction of the image of the second modality type itself. Since in the reconstruction process during training, both the input and output data are images of the second modality type, the problem of pixel mismatch is avoided. The features of the image of the first modality type and the image of the second modality type are supervised and constrained by the L2 norm, thereby indirectly realizing the style transfer from the image of the first modality type to the image of the second modality type, and obtaining a feature distribution with semantic consistency during this process. Thus, it is possible to convert an image of the first modality type into an image of the first modality type with a different modality type with high quality;

[0068] b. The cross-modal medical image style transfer method based on the feature space diffusion model of the present invention can adjust the operations of the feature space diffusion model during the noise addition and denoising processes based on the dynamic similarity mask, and show the characteristics of the feature distributions corresponding to the images of the two modality types. Taking the non-contrast CT image and the contrast-enhanced CT image as examples, the similar regions correspond to the relatively similar anatomical structure information in the two-phase CT data, while the non-similar regions correspond to the regions in the contrast-enhanced CT image that are different from the non-contrast CT image due to the addition of contrast agent; as the iteration of model training progresses, the focus of the feature space diffusion model dynamically shifts from the large structure similar regions with rich semantics to the contrast-enhanced regions with obvious contrast, thereby further improving the generation performance and generation efficiency of the present invention. In addition, compared with hard conditions such as segmentation results and text prompts, the similarity mask obtained through the model posterior is easier to obtain in practical applications and can be more directly and effectively incorporated into the model training process. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0070] Figure 1 It is a schematic diagram of the architecture of the first cross-modal style transfer model provided for an exemplary embodiment of the present invention;

[0071] Figure 2Schematic diagram of the training process of the first cross-modal style transfer model provided by an exemplary embodiment of the present invention;

[0072] Figure 3 Schematic diagram of the application process of the first cross-modal style transfer model provided by an exemplary embodiment of the present invention;

[0073] Figure 4 Schematic diagram of the process of the first cross-modal image style transfer method provided by an exemplary embodiment of the present invention;

[0074] Figure 5 Schematic diagram of the architecture of the feature space diffusion model provided by an exemplary embodiment of the present invention;

[0075] Figure 6 Schematic diagram of the training process of the feature space diffusion model with features as input provided by an exemplary embodiment of the present invention;

[0076] Figure 7 Schematic diagram of the application process of the feature space diffusion model provided by an exemplary embodiment of the present invention;

[0077] Figure 8 Schematic diagram of the training process of the feature space diffusion model with images as input provided by an exemplary embodiment of the present invention;

[0078] Figure 9 Schematic diagram of the training process of the pixel space reconstruction model provided by an exemplary embodiment of the present invention;

[0079] Figure 10 Schematic diagram of the application process of the second cross-modal style transfer model based on the feature space diffusion model provided by an exemplary embodiment of the present invention. Detailed implementation manners

[0080] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0081] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0082] Based on the prior art, the present application has found that the prior art has at least the following deficiencies:

[0083] I. Insufficient accuracy in pixel alignment of cross-modal images: Due to the patient's respiratory movement and the changes of various tissues or organs of the patient at different time points, the registration of the plain scan CT and the enhanced scan CT in the pixel space is often not ideal. The mainstream generation models (such as Pix2pix) usually assume that the input and output images are already precisely aligned and rely on this premise to learn the complex mapping relationship between paired images. However, this assumption may limit the generalization ability of the model in practical applications [8] , especially in the field of medical images, where the requirement for the accuracy of anatomical structures is higher, this limitation is particularly obvious;

[0084] II. The guidance in the feature space cannot balance the accuracy of the generated image and the focused target area: The current image generation models can be roughly divided into two technical routes: conditional guidance [9] and unconditional guidance.

[10] The main difference between them lies in whether conditional information is introduced into the model training process through an additional encoding mechanism. Conditional guidance models usually encode specific conditions (such as the segmentation result or text description of the region of interest in the image) as feature vectors and fuse them with the original image features, thereby enhancing the model's learning of these specific regions or contents. This way can make the model more focused on the key regions, so as to more accurately show the target features in the generated results. Although unconditional guidance models can be independent of certain specific conditions, due to the lack of targeted guidance, they lack accuracy in the process of generating specific regions. In actual situations, the supplement of specific conditions is often lacking in the generation process, which makes it difficult for the model to focus on the anatomical structures or lesion regions of interest. The references [8] to

[10] cited above are as follows:

[0085] [8]Kong L,Lian C,Huang D,et al.Breaking the dilemma of medical image-to-image translation[J].Advances in Neural Information Processing Systems,2021,34:1964-1978.

[0086] [9]Li Y,Shao H C,Liang X,et al.Zero-shot medical image translationvia frequency-guided diffusion models[J].IEEE transactions on medicalimaging,2023.

[0087]

[10] M,Dalmaz O,Dar S U H,et al.Unsupervised medical imagetranslation with adversarial diffusion models[J].IEEE Transactions on MedicalImaging,2023.

[0088] Based on the deficiencies of the above-mentioned existing technologies and combined with the analysis of practical applications and model accuracy, the present invention focuses on reducing the model error range caused by unaligned paired data and optimizing the technical process of unconditional guidance. Based on this, on the basis of the existing mainstream models, the present invention proposes an unconditional guidance model applicable to cross-modal generation of medical images, which can better meet the actual needs in the process of medical image generation.

[0089] In an embodiment of the present invention, a cross-modal image style transfer method is provided for converting an image of a first modal type into an image of a second modal type. Refer to Figure 1 and Figure 2 , and the method includes the following steps:

[0090] Obtain an initial image of a first modal type of an object, and input the initial image into a pre-constructed cross-modal style transfer model;

[0091] Wherein, the cross-modal style transfer model is pre-established through the following steps:

[0092] Collect a learning sample set, and each sample in the learning sample set includes an image of a first modal type and an image of a second modal type of the same object;

[0093] Design a basic model, where the basic model includes a first encoder and a second encoder; wherein, the first encoder and the second encoder are configured to extract features from an image;

[0094] Use the learning sample set to optimize and train the basic model in the following way:

[0095] Input the first-modal-type image in the sample into the first encoder to obtain a first image feature; and input the second-modal-type image in the sample into the second encoder to obtain a second image feature;

[0096] Use a preset total loss function to train the basic model so that the first image feature extracted by the first encoder of the trained model aligns with the second image feature extracted by the second encoder, obtaining the cross-modal style transfer model; the total loss function is at least related to a similarity loss, and the similarity loss is determined according to the first image feature and the second image feature;

[0097] The cross-modal style transfer model uses the first encoder therein to extract target image features from the initial image and reconstructs a target image of the second modal type based on the target image features.

[0098] Following clinical norms, clinicians tend to use non-contrast CT rather than contrast-enhanced CT. However, the clarity of non-contrast CT in visualizing anatomical structures is often insufficient. To solve this problem, many generative models are trained one-to-one based on paired non-contrast and contrast-enhanced CT data to improve the clarity of anatomical structures in non-contrast CT. However, most studies suffer from model accuracy problems caused by misaligned images and insufficient utilization of cross-modal data.

[0099] Therefore, based on the theoretical basis of representation learning, the present invention designs a generative structure with dual encoders and a multi-combination loss function to obtain a better feature representation. As Figure 1 shown, for paired non-contrast CT (x n ) and contrast-enhanced CT (x c ), the present invention simultaneously obtains the features of both through dual encoders, namely the feature of non-contrast CT (E n ) and the feature of contrast-enhanced CT (E c ). In this process, the dual encoder is composed of two encoders with the same structure, and the alignment of paired input images in the feature space is achieved through a similarity loss (L sim ). It should be noted that, in order to achieve a better alignment effect, a feature optimizer is added at the encoder end of the non-contrast CT input to obtain a better feature representation of the non-contrast image (Pn ).

[0100] In one embodiment of the present invention, the basic model further includes a feature optimizer, a quantization module, a decoder and a classifier. The feature optimizer is configured to optimize the features of its input to obtain optimized features. Specifically, the feature optimizer is mainly composed of a 3d convolution layer (Conv3d), a normalization layer (Group Norm) and an activation function layer (Sigmoid) connected in sequence, and obtains information with more structural characteristics by performing deeper feature extraction on the features of its input.

[0101] The quantization module is configured to quantize the features of its input to obtain quantized features. The decoder is configured to decode the image features of its input to obtain a reconstructed image; the classifier is configured to penalize the areas of the reconstructed image with poor reconstruction effect to improve the quality of the reconstructed image. Generally speaking, the adversarial generative network is mainly composed of a generator and a classifier, and the two achieve high-quality image generation in an adversarial manner. Based on the structure of the adversarial generative network, the present invention also designs a classifier (Discriminator) to identify the reconstructed image, and uses the classification loss (L dis ) Penalizes areas with poor reconstruction results, thereby improving the quality of the reconstructed image.

[0102] In one embodiment of the present invention, Figure 1 , Figure 2 and Figure 4 As shown, the process of optimizing the training of the basic model also includes the following steps.

[0103] The first image feature is quantized by using the quantization module to obtain a quantized feature; the total loss function is also at least related to the quantization loss, and the quantization loss is determined according to the first image feature and the quantized feature.

[0104] Regarding the quantization loss, preferably, the second image feature and the quantized feature after quantization are respectively calculated based on the square of the L2 norm. The weighted calculation is performed, and the calculation formula of the quantization loss is expressed as follows:

[0105]

[0106] Among them, L quan represents the quantization loss, E c represents the second image feature, detach() means pausing the gradient return of training during model training. represents the quantitative feature.

[0107] The decoder is used to reconstruct an image of a second modality type based on the second image feature to obtain a reconstructed image of the second modality type; the total loss function is further at least related to a reconstruction loss, and the reconstruction loss is based on the second modality type image and the reconstructed image of the second modality type.

[0108] Regarding the reconstruction loss L rec , preferably, the similarity between the enhanced-phase CT (x c ) and the enhanced-phase CT after reconstruction is calculated based on the square of the L1 norm, and the calculation formula of the reconstruction loss is expressed as follows: The calculation formula of the reconstruction loss is as follows:

[0109]

[0110] where L rec represents the reconstruction loss, x c represents the second modality type image, represents the reconstructed image of the second modality type.

[0111] The classifier is configured to determine a classification loss, and the classification loss is determined based on the second image and the reconstructed image of the second modality type reconstructed based on the second image feature; the total loss function is further at least related to the classification loss.

[0112] Regarding the classification loss L dis , preferably, the second modality type image and the reconstructed image of the second modality type are calculated through a binary cross-entropy loss function, and the calculation formula of the classification loss is expressed as follows:

[0113]

[0114] where L dis represents the classification loss, D(·) represents the classifier, x c represents the second modality type image, represents the reconstructed image of the second modality type.

[0115] In this embodiment, in the process of determining the similarity loss according to the first image feature and the second image feature, the feature optimizer is first used to optimize the first image feature extracted from the image by the first encoder to obtain an optimized feature. Then, the similarity loss is calculated based on the optimized feature and the second image feature.

[0116] Regarding the similarity loss L sim , preferably, the optimized feature (P n ) after optimization and the second image feature (E c ) are calculated through the L2 norm, and its calculation formula is as follows:

[0117] L sim = ||E c - P n || 2

[0118] Among them, L sim represents the similarity loss, E c represents the second image feature, and P n represents the optimized feature.

[0119] The loss functions involved in each of the above parts are all involved in the training process of the autoencoder (AutoEncoder). That is, during the training process, the quantization loss, reconstruction loss, classification loss, and similarity loss are respectively less than their respective preset thresholds to make their corresponding modules meet the training requirements.

[0120] Regarding the unified paradigm of each loss, the present invention designs the following total loss function based on VQ - GAN, and the calculation method of the total loss function is expressed as follows:

[0121] L total = α r L rec + α q L quan + α s L sim + α c L dis

[0122] Among them, α r , α q , α s , α d respectively represent the preset weight values of the reconstruction loss, quantization loss, similarity loss, and classification loss.

[0123] By designing the above - combined loss function, the autoencoder can not only generate high - quality images but also effectively align the features of cross - modal data (plain - scan CT and enhanced - scan CT). This feature alignment provides an optimization space for denoising based on the diffusion model in the feature space, thereby further improving the generation performance and application effect of the present invention in converting the first - modality type of image into the second - modality type of image.

[0124] In this embodiment, when the total loss function is less than the preset threshold, the training of the model is completed, and the trained model is the cross - modal style transfer model. The step - by - step process of using the cross - modal style transfer model to convert the first - modality type of image into the second - modality type of image is as Figure 3 shown.

[0125] Obtain an initial image of the first modality type of an object, and input the initial image into a pre-constructed cross-modal style transfer model. The cross-modal style transfer model uses a first encoder therein to extract target image features from the initial image and output a feature optimizer; the feature optimizer optimizes the target image features to obtain optimized features and outputs a quantization module; the quantization module performs quantization processing on the target image features to obtain target quantization features and outputs them to an adversarial generation network module composed of a decoder and a classifier, and the decoder decodes and reconstructs the target quantization features to obtain a target image of the second modality type.

[0126] In another embodiment of the present invention, different from the above embodiment is the design of the total loss function. In this embodiment, the total loss function is not determined according to the quantization loss, reconstruction loss, classification loss, and similarity loss, but is determined according to a part of the quantization loss, reconstruction loss, classification loss, and similarity loss.

[0127] In another embodiment of the present invention, the difference from the above embodiment lies in the optimization featureizer. In this embodiment, the feature optimizer is not configured at the output end of the first encoder. Regarding the similarity loss L sim , it is directly calculated according to the first image feature and the image feature, and its calculation formula is expressed as follows:

[0128] L sim =||E c -E n || 2

[0129] wherein, L sim represents the similarity loss, E c represents the second image feature, and E n represents the first image feature.

[0130] The cross-modal image style transfer method provided by the above embodiments, in the process of converting an image of the first modality type into an image of the second modality type, realizes the alignment of cross-modal medical images at the feature level through representation learning, thereby effectively reducing the problems caused by the misalignment of pixels and anatomical structures in cross-modal data in practical applications.

[0131] In practical applications, taking non-contrast CT images and contrast-enhanced CT images as an example, that is, the non-contrast CT images are the first images and the contrast-enhanced CT images are the second images, the above cross-modal image style transfer method cleverly decomposes the style transfer task from non-contrast CT to contrast-enhanced CT into two subtasks: one is to achieve the alignment of non-contrast CT and contrast-enhanced CT at the feature level ( Figure 2The process indicated by the two dashed lines in the lower middle part); second, the reconstruction of the enhanced-phase CT itself ( Figure 2 The process indicated by the semi-dashed line in the upper middle part). Since both the input and output data are enhanced-phase CTs during the reconstruction process, the problem of pixel mismatch is avoided. At the same time, the features of the plain scan-phase CT and the enhanced-phase CT are supervised and constrained by the L2 norm, thereby indirectly realizing the style transfer from the plain scan-phase CT to the enhanced-phase CT, and obtaining a feature distribution with semantic consistency during this process. Thus, it is possible to convert the plain scan-phase CT image into an enhanced-phase CT with a different modal type with high quality.

[0132] The above-mentioned cross-modal image style transfer method is a technical solution of the present invention for converting an image of a first modal type into an image of a second modal type based on the Pixel Space. Through the design of the model structure and the constraint of the loss function, the feature alignment of cross-modal data is achieved. However, using an adversarial generative model as the generative model has certain instability in practical applications and also has a high demand for computing resources. Therefore, the present invention introduces a diffusion model in the Latent Space, and optimizes the feature distribution (L diff Calculation formula) to construct a cross-modal style transfer model with more stable performance and higher computational efficiency.

[0133] In an embodiment of the present invention, a cross-modal image style transfer method is provided. Refer to Figures 5 to 7 The cross-modal image style transfer method further includes the following steps:

[0134] Using a pre-constructed feature space diffusion model to convert the target image feature into a target fusion feature; the cross-modal style transfer model reconstructs a target image of the second modal type based on the target fusion feature;

[0135] Among them, the feature space diffusion model is pre-established through the following steps:

[0136] Collect a diffusion learning sample set, and each sample in the diffusion learning sample set includes the first image feature and the second image feature;

[0137] Design a basic diffusion model, the basic diffusion model includes a dynamic similarity mask module and a diffusion module, the dynamic similarity mask module is configured to obtain a mask of its input image feature, and the diffusion module is configured to perform a masking operation on the first image feature using the mask to obtain a fused feature;

[0138] Using the diffusion learning sample set, optimize and train the basic diffusion model in the following manner:

[0139] Input the first image feature and the second image feature into the dynamic similarity mask module to obtain a dynamic similarity mask for the first image feature and the second image feature;

[0140] Input the first image feature and the dynamic similarity mask into the diffusion module to obtain a first fused feature;

[0141] And train the basic diffusion model using a preset diffusion loss function, so that the first fused feature output by the trained model incorporates the structural information possessed by the second image feature on the basis of the first image feature, obtaining the feature space diffusion model; the preset diffusion loss function is at least related to the dynamic similarity mask;

[0142] The feature space diffusion model uses the diffusion module therein to incorporate the relevant structural information possessed by the image features of the second modality type of image on the basis of the target image feature to obtain a first target fused feature, and determines the target fused feature according to the first target fused feature.

[0143] In an embodiment of the present invention, the first target fused feature is determined as the target fused feature. In other better embodiments of the present invention, introducing a noise addition process and a denoising process for the first target fused feature can obtain a target fused feature with better quality.

[0144] In an embodiment of the present invention, the basic diffusion model further includes a noise processing module. As Figure 5 shown, the noise processing module includes a noise addition module and a denoising module, wherein the noise addition module is configured to perform noise addition processing on its input; the denoising module is configured to perform denoising processing on its input.

[0145] The present invention calculates the similarity matrix of the first image feature (preferably the optimized feature) and the second image feature through cosine similarity, and uses it as additional information to strengthen the original input feature of the model, that is, the first image feature E n or the optimized feature P n . In order to better utilize semantic information during the noise addition and denoising processes and guide the model to actively learn sparse distribution and low similarity regions in plain scan CT and enhanced scan CT, the present invention achieves this goal through the change of the dynamic similarity mask (Dynamic Similarity Mask, DSM). The calculation formula of the dynamic similarity mask DSM is as follows:

[0146]

[0147] Among them, DSM represents the dynamic similarity mask, τ represents the current training cycle number, Total epochs represents the total number of training cycles, <·,·> represents the cosine similarity between the first image feature and the second image feature, and the value range of <·,·> is [-1, 1]. is the cosine similarity matrix between the first image feature and the second image feature, which is used to change the value range of the similarity to [0, 1], and min[·,·] represents taking the minimum value between two elements. In this formula, by comparing with 1 and taking the minimum value, α is a scaling factor, which is a preset value, and α is used to control the change rate of the cycle ratio.

[0148] During the training process of the basic diffusion model, the present invention combines the dynamic similarity mask and the denoising process by designing a diffusion loss function (L diff ). The calculation formula of the diffusion loss function is as follows:

[0149]

[0150] Among them, DSM represents the dynamic similarity mask, E n represents the first image feature, and E represents the second image feature. represents the dot product operation. represents the random noise ∈ sampled from the standard normal distribution (N(0, 1)) at each time step t, combined with the input feature E n and E c to perform multiple samplings and calculate the error expectation between the output noise ∈ θ obtained by the model and the input noise ∈. ∈ represents the input noise, and the input noise is determined by the first image feature and the second image feature. ∈ θ represents the output noise, and the output noise is determined by the second fusion feature; represents the constraint by the square of the L2 norm. The time step t is a value uniformly randomly drawn from the set {1,…, T}, and T represents the total number of time steps in the denoising process.

[0151] Optimize and train the basic diffusion model using the above diffusion loss function to obtain the trained model, and use the trained model as the feature space diffusion model. Take the target image feature or the target optimization feature as the input of the feature space diffusion model. At this time, one input of the dynamic similarity mask module is the target image feature or the target optimization feature. Since there is no input item corresponding to the second image feature in the training process, the other input is 0, and the dynamic similarity mask module is a matrix of all 1s. Based on this, the trained diffusion module adds the image features of the second modality type of image on the basis of the target image feature or the target optimization feature, thereby obtaining the first target fusion feature, and further performing noise addition processing on the first target fusion feature to obtain the second fusion feature after adding Gaussian noise, and then performing step-by-step denoising processing with a total number of steps of T on the second fusion feature to obtain the third fusion feature, and the third fusion feature is used as the target fusion feature to be obtained.

[0152] By introducing the feature space diffusion model into the cross-modal style transfer model, using the target fusion feature as the input of the quantization module, and sequentially performing quantization processing and decoding reconstruction on the target fusion feature to obtain the target image of the second modality type.

[0153] In an embodiment of the present invention, a cross-modal medical image style transfer method is provided. The method is used to incorporate the differential structure information that the second modality type of image has and the first modality type of image does not have into the first modality type of image. Refer to Figure 8 , in this embodiment, the method includes the following steps:

[0154] Obtain the initial image of the first modality type of an object, and input the initial image into the pre-constructed feature space diffusion model;

[0155] Among them, the feature space diffusion model is pre-established through the following steps:

[0156] Collect a learning sample set, and each sample in the learning sample set includes the first modality type image and the second modality type image of the same object;

[0157] Design a basic diffusion model, which includes an encoder module, a dynamic similarity mask module, and a diffusion module; among them, the encoder module is configured to extract features from the image;

[0158] Use the learning sample set to optimize and train the basic model in the following manner:

[0159] Input the first-modal type image in the sample into the encoder module to obtain the first image feature; and input the second-modal type image in the sample into the encoder module to obtain the second image feature;

[0160] Input the first image feature and the second image feature into the dynamic similarity mask module to obtain a dynamic similarity mask for the first image feature and the second image feature;

[0161] Input the first image feature and the dynamic similarity mask into the diffusion module to obtain a first fused feature;

[0162] And train the basic diffusion model using a preset diffusion loss function, so that the first fused feature output by the trained model incorporates, on the basis of the first image feature, the structural information possessed by the second image feature that is different from the structural information possessed by the first image feature, to obtain the feature space diffusion model; the diffusion loss function is at least related to the dynamic similarity mask;

[0163] The feature space diffusion model uses the encoder module to extract features from the initial image to obtain target image features, and uses the diffusion module to incorporate, on the basis of the target image features, the structural information possessed by the image features of the second-modal type image to obtain a first target fused feature.

[0164] Preferably, the basic diffusion model further includes a noise processing module. As Figure 8 shown, the noise processing module is configured to perform noise addition processing and / or denoising processing on the first fused feature to obtain a second fused feature; the diffusion loss function is further at least related to the input noise and output noise of the basic diffusion model, the input noise is determined by the first image feature and the second image feature, and the output noise is determined by the second fused feature;

[0165] The feature space diffusion model uses the noise processing module therein to perform noise addition and / or denoising processing on the first target fused feature to obtain a second target fused feature; and determines the second target fused feature as the target fused feature.

[0166] Preferably, the basic diffusion model further includes a feature optimizer, and the feature optimizer is configured to perform optimization processing on the input image features to obtain optimized features, and the optimized features are used as the input of the noise processing module.

[0167] Among them, the diffusion loss L diff The calculation formula is as follows:

[0168]

[0169] Among them, DSM represents the dynamic similarity mask, and its calculation method is as described in the above embodiments and will not be elaborated here. E n represents the first image feature, represents the dot product operation, represents the random noise ∈ sampled from the standard normal distribution (N(0, 1)) at each time step t, combined with the input data x n and x c Perform multiple samplings, calculate the noise ∈ of the model output θ and the expected error between the real noise ∈, where ∈ represents the input noise, ∈ θ represents the output noise, represents being constrained by the square of the L2 norm, P n is the optimization feature, and the time step t is a value randomly and uniformly drawn from the set {1, …, T}, where T represents the total number of time steps in the denoising process.

[0170] Among them, the encoder module can be set in two ways. In one setting, the encoder module is a dual-encoder structure, which includes a first encoder and a second encoder. The first encoder is configured to extract the features of the first-modal type image to obtain the first image feature, and the second encoder is configured to extract the features of the second-modal type image to obtain the second image feature. In one setting, the encoder module is a single-encoder structure, which only includes one encoder. The encoder is used to extract the features of the first-modal type image and the second-modal type image respectively, and the first image feature and the second image feature are distinguished by the labels respectively configured for the first-modal type image and the second-modal type image.

[0171] In this embodiment, the cross-modal medical image style transfer method further includes converting the first-modal type image into the second-modal type image by using a pre-established pixel space reconstruction model.

[0172] Preferably, referring to Figure 9 , the basic model of the pixel space reconstruction model in this embodiment includes a feature encoder module, a feature optimizer, a quantization module, a decoder, and a classifier connected in sequence. In this embodiment, the functions of the feature optimizer, the noise processing module, the quantization module, the decoder, and the classifier are the same as those described in the above embodiments and will not be elaborated here.

[0173] Using the first-modal type image and / or the second-modal type image as the learning sample set, optimize and train the basic model of the pixel space reconstruction model by using the pixel space reconstruction loss function to obtain the trained model as the pixel space reconstruction model. In this embodiment, the calculation method of the pixel space reconstruction loss function is expressed as follows:

[0174] L′ total = k r L rec + k q L quan + k d L dis

[0175] Among them, L′ total represents the pixel space reconstruction loss function, L rec represents the reconstruction loss, L quan represents the quantization loss, L dis represents the classification loss, k r and k q and k d respectively represent the preset weight values of the reconstruction loss, quantization loss, diffusion loss, and classification loss.

[0176] It should be noted that the calculation methods of the reconstruction loss, quantization loss, and classification loss can adopt the calculation methods described in the above embodiments, and will not be elaborated here.

[0177] See Figure 10 , add the trained diffusion model to the trained pixel space reconstruction model. Specifically, merge the feature encoder module and feature optimizer in the diffusion model and the pixel space reconstruction model, use the output of the diffusion model as the input of the quantization module in the pixel space reconstruction model, and the quantization module and decoder in the pixel space reconstruction model perform quantization processing and reconstruction on the output (the first target fusion feature / second target fusion feature) of the diffusion model to obtain the target image of the second modality type.

[0178] The cross-modal medical image style transfer method based on the feature space diffusion model provided by the present invention can adjust the operations of the feature space diffusion model during the noise addition and denoising processes based on the dynamic similarity mask, and show the characteristics of the feature distributions corresponding to the images of the two modality types. Taking the plain scan CT image and the enhanced CT image as examples, the similar regions correspond to the relatively similar anatomical structure information in the two-phase CT data, while the non-similar regions correspond to the regions in the enhanced CT image that are different from the plain scan CT image due to the addition of contrast agent. As the model training iterates, the focus of the feature space diffusion model dynamically shifts from the large structure similar regions with rich semantics to the enhanced imaging regions with obvious contrast, thereby further improving the generation performance and generation efficiency of the present invention. In addition, compared with hard conditions such as segmentation results and text prompts, the similarity mask obtained through the model posterior is easier to obtain in practical applications and can be more directly and effectively incorporated into the model training process.

[0179] In an embodiment of the present invention, a cross-modal image style transfer system is provided. The cross-modal image style transfer system uses the cross-modal image style transfer method described in any of the above embodiments to convert an image of a first modal type into an image of a second modal type.

[0180] In an embodiment of the present invention, a computer-readable storage medium is provided for storing program instructions, and the program instructions are configured to be called to execute the steps of the method described in any of the above embodiments.

[0181] It should be noted that the above embodiments of the cross-modal image style transfer system and the computer-readable storage medium belong to the same inventive concept as the embodiments of the cross-modal image style transfer method. By reference, all the contents of the embodiments of the cross-modal image style transfer method are incorporated into the embodiments of the cross-modal image style transfer system and the computer-readable storage medium.

[0182] It should be noted that in this article, relational terms such as are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, article or device comprising the element.

[0183] The above are only specific embodiments of the present application. It should be pointed out that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A cross-modal image style transfer method for converting an image of a first modality type into an image of a second modality type, characterized in that: The method comprises the following steps: Acquire an initial image of a first modality type of an object, and input the initial image into a pre-built cross-modal style transfer model; The cross-modal style transfer model is established in advance through the following steps: Collecting a learning sample set, each sample in the learning sample set includes a first modality type image and a second modality type image of the same object; Designing a basic model, the basic model comprising a first encoder and a second encoder; wherein the first encoder and the second encoder are configured to extract features from an image; Using the learning sample set, the basic model is optimized and trained in the following manner: Inputting a first modality type image in the sample into the first encoder to obtain a first image feature; and inputting a second modality type image in the sample into the second encoder to obtain a second image feature; The basic model is trained using a preset total loss function so that a first image feature extracted by a first encoder of the trained model is aligned with a second image feature extracted by a second encoder, thereby obtaining the cross-modal style transfer model; the total loss function is at least related to a similarity loss, and the similarity loss is determined according to the first image feature and the second image feature; The cross-modal style transfer model uses a first encoder therein to extract target image features from the initial image, and reconstructs a target image of a second modality type based on the target image features.

2. The cross-modal image style transfer method according to claim 1, characterized in that: The basic model further includes a quantization module. During the basic model optimization training process, the quantization module is configured to perform quantization processing on the first image feature to obtain a quantized feature; The total loss function is also related to at least a quantization loss, and the quantization loss is determined according to the first image feature and the quantization feature; The cross-modal style transfer model uses a quantization module therein to quantize the target image features to obtain target quantized features, and reconstructs a target image of the second modality type based on the target quantized features.

3. The cross-modal image style transfer method according to claim 2, characterized in that: The calculation formula of the quantization loss is expressed as follows: Among them, L quan represents the quantization loss, E c represents the second image feature, detach() means pausing the gradient return of training during model training. represents the quantitative feature.

4. The cross-modal image style transfer method according to claim 1, characterized in that: The base model further includes a decoder, and during the base model optimization training process, the decoder is configured to reconstruct an image of a second modality type based on the second image feature to obtain a reconstructed image of the second modality type; the total loss function is also at least related to a reconstruction loss, and the reconstruction loss is based on the second modality type image and the second modality type reconstructed image; The cross-modal style transfer model utilizes a decoder therein to reconstruct a target image of a second modality type based on the target image features.

5. The cross-modal image style transfer method according to claim 4, characterized in that: The calculation formula of the reconstruction loss is expressed as follows: Among them, L rec represents the reconstruction loss, x c represents the second modality type image, Represents the second modality type reconstructed image.

6. The cross-modal image style transfer method according to claim 1, characterized in that: The basic model also includes a classifier. During the basic model optimization training process, the classifier is configured to determine a classification loss, and the classification loss is determined based on the second image and a reconstructed image of a second modality type reconstructed based on features of the second image; the total loss function is also at least related to the classification loss.

7. The cross-modal image style transfer method according to claim 6, characterized in that: The calculation formula of the classification loss is expressed as follows: Among them, L dis represents the classification loss, D(·) represents the classifier, and x c represents the second modality type image, Represents the second modality type reconstructed image.

8. The cross-modal image style transfer method according to claim 1, characterized in that: The similarity loss is determined according to the first image feature and the second image feature, and the calculation formula of the similarity loss is expressed as follows: THE sim =||And n -AND n ||2 Among them, L sim represents the similarity loss, E c represents the second image feature, E n represents the first image feature.

9. The cross-modal image style transfer method according to claim 1, characterized in that: The base model further includes a feature optimizer configured to optimize features extracted from the image by the first encoder; The cross-modal style transfer model optimizes the target image features using the feature optimizer therein to obtain optimized features, and reconstructs a target image of the second modality type based on the optimized features.

10. The cross-modal image style transfer method according to claim 9, characterized in that: The feature optimizer includes a 3D convolution layer, a normalization layer and an activation function layer connected in sequence.

11. The cross-modal image style transfer method according to claim 1, characterized in that: The method further includes converting the target image features into target fusion features using a pre-constructed feature space diffusion model; reconstructing the target image of the second modality type based on the target fusion features using the cross-modal style transfer model; The feature space diffusion model is established in advance through the following steps: Collecting a diffusion learning sample set, each sample in the diffusion learning sample set includes the first image feature and the second image feature; Designing a basic diffusion model, the basic diffusion model includes a dynamic similarity mask module and a diffusion module, the dynamic similarity mask module is configured to obtain a mask of its input image features; Using the diffusion learning sample set, the basic diffusion model is optimized and trained in the following manner: Inputting the first image feature and the second image feature into the dynamic similarity mask module to obtain a dynamic similarity mask about the first image feature and the second image feature; Inputting the first image feature and the dynamic similarity mask into the diffusion module to obtain a first fusion feature; The basic diffusion model is trained using a preset diffusion loss function, so that a first fusion feature output by the trained model incorporates structural information possessed by the second image feature on the basis of the first image feature, thereby obtaining the feature space diffusion model; the preset diffusion loss function is at least related to the dynamic similarity mask; The feature space diffusion model utilizes the diffusion module therein to integrate the structural information possessed by the image features of the second modality type of image on the basis of the target image features to obtain the first target fusion feature, and determines the target fusion feature based on the first target fusion feature.

12. The cross-modal image style transfer method according to claim 11, characterized in that: The basic diffusion model further includes a noise processing module, which is configured to perform noise addition and / or denoising on the first fused feature to obtain a second fused feature; the diffusion loss function is also at least related to input noise and output noise of the basic diffusion model, the input noise is determined by the first image feature and the second image feature, and the output noise is determined by the second fused feature; The feature space diffusion model uses the noise processing module therein to perform noise addition and / or denoising processing on the first target fusion feature to obtain a second target fusion feature; And determine the second target fusion feature as the target fusion feature.

13. The cross-modal image style transfer method according to claim 12, characterized in that: The diffusion loss function is expressed by the following formula: Wherein, DSM represents the dynamic similarity mask, E n represents the first image feature, E represents the second image feature, represents the dot product operation, represents the random noise ∈ sampled from the standard normal distribution (N(0,1)) at each time step t, combined with the input data E n and E c Perform multiple sampling and output noise ∈ θ The expected error between the input noise ∈, ∈ represents the input noise, ∈ θ represents the output noise, Represents the square of the L2 norm.

14. The cross-modal image style transfer method according to claim 11 or 12, characterized in that: The calculation formula of the dynamic similarity mask is as follows: Wherein, DSM represents the dynamic similarity mask, τ represents the current training cycle number, Total epochs represents the total training cycle number, <·,·> represents the cosine similarity between the first image feature and the second image feature, and the value range of <·,·> is [-1,1], The range of similarity is changed to [0,1], min[·,·] means taking the minimum value between two elements, and α is the scaling factor, which is the preset value.

15. The cross-modal image style transfer method according to claim 11, characterized in that: Also includes: The first target fusion feature is determined as the target fusion feature.

16. A cross-modal image style transfer system, characterized in that: The cross-modal image style transfer system converts an image of a first modality type into an image of a second modality type using the cross-modal image style transfer method as described in any one of claims 1 to 15.

17. A computer-readable storage medium for storing program instructions, characterized in that: The program instructions are configured to be called to execute the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and system for converting RGB image into NIR image

    CN117808669A

  • Unsupervised pedestrian re-identification method and system based on modal invariance modeling

    CN118015657A

  • A method for transferring the style of an image and an apparatus therefor

    KR102592666B1

  • Image Processing Method and Apparatus, Storage Medium and Electronic Device

    US20200202111A1

  • Method and apparatus for training machine learning model, apparatus for video style transfer

    US20210256304A1