Unsupervised image conversion imaging method based on Schrodinger bridge theory

By employing an image conversion method based on Schrödinger bridge theory, and using adversarial training of the time-conditional generator and discriminator, combined with multiple constraint mechanisms, the problems of training instability and insufficient generation quality in image conversion are solved, achieving efficient and high-quality conversion between complex image domains.

CN121353104APending Publication Date: 2026-01-16BEIHANG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511527535.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing image transformation methods struggle to balance data dependence, training stability, generation quality, and computational efficiency, especially when transforming complex image distributions, where problems such as training instability, loss of detail, or artifacts arise.

Method used

An image domain-to-domain transformation framework is constructed based on Schrödinger's bridge theory. A time-conditional generator is used to simulate the smooth evolution path of the image, and an adversarial training is formed by combining it with a discriminator. At the same time, salient content guidance constraints, global feature consistency constraints, and contrastive learning constraints are integrated, and the model is optimized through a composite loss function.

Benefits of technology

It achieves stable, high-quality, and efficient image transformation between complex image domains, improving the stability of model training and the fidelity of generated images. In particular, it can reliably preserve key information and details of the image when there are significant differences between the source and target domains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121353104A_ABST
    Figure CN121353104A_ABST
Patent Text Reader

Abstract

The invention discloses an unsupervised image conversion imaging method based on the Schrodinger bridge theory, and the method comprises the following steps: 1, constructing an image domain-domain conversion frame based on the Schrodinger bridge theory, simulating a smooth evolution path from a source domain image to a target domain image through a time condition generator, and forming antagonism training in combination with a discriminator; 2, integrating a saliency content guide constraint, a global feature consistency constraint and a contrast learning constraint in the conversion framework so as to enhance semantic content retention and detail generation capability; 3, establishing a composite loss function, and carrying out weighted combination on Schrodinger bridge path loss, adversarial loss and each auxiliary constraint loss for guiding model optimization; and 4, training a time condition generator and a discriminator through an optimization algorithm, and carrying out stable and high-fidelity target domain conversion on the source domain image by utilizing the generator after training is completed. According to the method, the stability of model training and the quality and fidelity of the generated image are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and deep learning, and in particular to an unsupervised image-to-image conversion method based on Schrödinger bridge theory. Background Technology

[0002] Image transfer technology, a core topic in computer vision and graphics, aims to transfer images from a source domain to a target domain while preserving the core content information of the source image and endowing it with the unique style, attributes, or modal features of the target domain. This technology has demonstrated broad application prospects and significant value in many fields.

[0003] A key application of image transformation is domain adaptation. Its core objective is to bridge the gap between different data domains caused by differences in imaging conditions, sensor characteristics, or data distribution, enabling knowledge acquired or models trained in one data domain to be effectively applied to another data domain with different characteristics. For example, in medical image analysis, converting multispectral or hyperspectral imaging data into pathological images that simulate traditional chemical staining is an important domain adaptation technique that helps achieve rapid and non-destructive histological analysis.

[0004] In recent years, the rapid development of deep learning, especially generative models, has greatly promoted the advancement of image translation technology. Current mainstream methods can be broadly categorized into two types: supervised and unsupervised. Supervised image translation methods, such as the Pix2Pix model based on conditional generative adversarial networks, typically require a large amount of precisely paired training data, meaning that each source domain image has a corresponding target domain image. While these methods can produce high-quality results, their strong dependence on paired data limits their application in many scenarios where paired samples are difficult to obtain.

[0005] To overcome data dependency bottlenecks, unsupervised image translation methods have emerged, requiring only domain-independent image sets for training. Representative methods include CycleGAN, which uses cycle consistency loss to constrain the translation process. Another type of method, CUT, borrows from contrastive learning, maximizing the mutual information between input and output image patches for unsupervised translation. Furthermore, diffusion models, as a powerful emerging generative model, have also been applied to image translation tasks due to their excellent generation quality and training stability, generating target domain images by guiding the diffusion process conditioned on the source image.

[0006] However, existing image transformation methods still face several challenges. Supervised methods are costly to acquire data. Unsupervised methods, such as CycleGAN, often encounter problems such as training instability and pattern collapse, and cycle consistency constraints may lead to loss of detail or artifacts. Contrastive learning-based methods still have room for improvement in the fineness of generated quality. Although diffusion models produce excellent results, their training and inference processes typically require huge computational resources and long time, limiting their application in resource-constrained or real-time-critical scenarios. Therefore, achieving a balance between data dependence, training stability, generated quality, and computational efficiency, especially in handling complex image distribution transformations to achieve stable, efficient, and high-quality transformations, remains a key technical challenge that urgently needs to be addressed in the field. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art and provide an unsupervised image conversion method based on Schrödinger bridge theory. Based on the conversion framework of Schrödinger bridge theory, multiple auxiliary constraint mechanisms are integrated to significantly improve the stability of model training and the quality and fidelity of generated images.

[0008] The objective of this invention is achieved through the following technical solution: an unsupervised image conversion method based on Schrödinger bridge theory, comprising the following steps:

[0009] Step 1: Construct an image domain-to-domain transformation framework based on Schrödinger bridge theory, use a temporal condition generator to simulate the smooth evolution path from the source domain image to the target domain image, and combine it with a discriminator to form adversarial training;

[0010] This step aims to establish a stochastic process that can simulate the smooth evolution of an image from a source domain distribution to a target domain distribution. A temporal conditional generator is employed to parameterize and learn the dynamics of this evolutionary path. Simultaneously, a discriminator network is introduced to evaluate the realism of the target domain image output by the generator, constituting an adversarial training mechanism.

[0011] Step 1 includes:

[0012] The time condition generator employs a U-shaped network architecture, which receives the current evolution time step. Image status and time step The image itself serves as input, and through learning, the image state at the next time step is predicted. By starting from the initial time By the final time Iterative calls at multiple discrete time steps enable the processing of images from the source domain. To the final target domain image The conversion.

[0013] Step 2: Integrate salient content guidance constraints, global feature consistency constraints, and contrastive learning constraints into the transformation framework to enhance semantic content preservation and detail generation capabilities;

[0014] This step addresses the issue that Schrödinger's Bridge theory itself does not directly include explicit content preservation or detail fidelity constraints. Building upon its basic framework, this step innovatively integrates multiple auxiliary constraint modules. These modules aim to guide the generator from different levels to better preserve key information from the source image and generate high-quality details that conform to the characteristics of the target domain during the transformation process.

[0015] Step 2 includes:

[0016] The module designs and integrates multiple auxiliary constraint mechanisms to enhance content and detail preservation, specifically as follows: A salient content-guided constraint module identifies key semantic regions in the source image (e.g., salient maps obtained through simple image intensity-based segmentation). This module first generates salient maps representing the key regions of both the source and generated images using a fixed grayscale threshold segmentation method. To ensure the differentiability of the training process, this segmentation operation is approximated using a smooth sigmoid function. Subsequently, the module calculates the mean squared error (MSE) between the two salient maps as a consistency loss. This is to enhance the model's ability to retain and accurately convert important content details.

[0017] Global Feature Consistency Constraint Module: Utilizes a deep convolutional neural network VGG16 on a large image dataset as a fixed feature extractor. It extracts activation feature maps from multiple intermediate layers of the network for both the source image and the final generated pseudo-target domain image, and calculates the mean squared error between these corresponding layer feature maps as the global feature consistency loss. This ensures that both maintain consistency in global semantic content and high-level structural features at the multi-scale abstraction level.

[0018] Step 3: Establish a composite loss function, which combines the Schrödinger bridge path loss, adversarial loss, and various auxiliary constraint losses in a weighted manner to guide model optimization;

[0019] This step combines the core Schrödinger bridge path loss and adversarial loss from Step 1 with the loss terms generated by the various auxiliary constraint modules in Step 2, forming a composite objective function. This composite loss function comprehensively guides the model in learning inter-domain mapping relationships while taking into account the realism, content consistency, structural integrity, and detail richness of the generated images.

[0020] A composite loss function is constructed to collaboratively optimize the model, as follows: The composite loss function is expressed as:

[0021]

[0022] To combat losses, a discriminator and generator Calculations show that These are the parameters for the generator and the discriminator, respectively. In the generator... for Intermediate images under time conditions, For the distribution of this state, For random noise sampled from a standard normal distribution, The noise distribution is used to increase the randomness of the generation process, and the same applies to subsequent operations. Adversarial loss is used to enhance the realism of the generated images.

[0023]

[0024] The Schrödinger bridge path loss is used to drive the model to learn the optimal random path from the source domain distribution to the target domain distribution, and includes a fidelity term and an entropy regularization term;

[0025]

[0026] For the target domain image, For entropy regularization function, For the inverse conditional probability learned by the model, , Weights for time steps

[0027] To contrast the learning loss, it is used to enhance the texture realism and detail diversity of local regions in the generated image. It learns fine-grained correspondences by maximizing the mutual information between the features of a patch in the generated image and its corresponding ground truth patch features in the source / target domain, while minimizing the mutual information between the patch and its non-corresponding features. From the source image Middle position Extracted query features. Generating a target domain image from a source domain image same position Extracted positive sample key features. These are negative sample key features extracted from other locations or images. The cosine similarity function is used. For temperature coefficient, For the set of layers used for comparison, For the first The set of spatial locations of a layer.

[0028]

[0029] The saliency consistency loss is used to guide the model to accurately identify and preserve semantically important regions or structures in the source image during image transformation. It ensures the correct transfer and transformation of key content by comparing the saliency maps of the source image and the corresponding regions in the generated image; where... and Image samples from the source domain and the generated target domain, respectively.

[0030]

[0031] Where G is the generator from the source domain to the target domain. , Foreground extraction operator based on grayscale threshold, The corresponding threshold is in the form of:

[0032]

[0033] A global feature consistency loss is used to ensure that the generated image maintains consistency with the source image in terms of global semantic content and high-level structural features at multiple scale abstraction levels. Deep features extracted by the deep convolutional network VGGnet are compared to capture content and structural information beyond pixel-level similarity. For the source domain image, To generate an image from the source domain image to the target domain image, This indicates that the VGG16 network is in the... Feature maps output by the layer It is the first The dimensions of the layer feature map (channels, height, width). The set representing the number of network layers

[0034]

[0035] Step 4: Train the temporal condition generator and discriminator by optimizing the algorithm, and after training is completed, use the generator to perform stable and high-fidelity target domain transformation on the source domain image.

[0036] This step employs an optimization algorithm (such as gradient descent) to minimize the constructed composite loss function, jointly training the temporal conditional generator and discriminator. After training, the resulting temporal conditional generator can accept source domain images as input and stably convert them into high-quality, high-fidelity target domain images.

[0037] The beneficial effects of this invention are: by innovatively integrating multiple auxiliary constraint mechanisms on the basis of the Schrödinger bridge theory conversion framework, this invention aims to significantly improve the stability of model training, the quality and fidelity of generated images, especially in application scenarios where there are significant differences between the source and target domains or where fine structure and texture preservation is required, it can achieve more reliable and high-quality image conversion. Attached Figure Description

[0038] Figure 1 This is a flowchart of the method of the present invention;

[0039] Figure 2 This is an overall architecture diagram of the image conversion model of the present invention;

[0040] Figure 3 This is a schematic diagram of the network structure of the time condition generator;

[0041] Figure 4 This is a flowchart of the image conversion method in the embodiment. Detailed Implementation

[0042] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0043] This invention considers the Schrödinger bridge theory, which offers a promising new framework for transformations between complex data distributions. Originating in physics, this theory aims to find the optimal stochastic evolutionary path connecting two given probability distributions. Its direct modeling of the dynamic migration process between arbitrary initial and final distributions makes it naturally suitable for transformations between image domains. By parameterizing and learning this optimal evolutionary path, the Schrödinger bridge-based method is expected to overcome the shortcomings of existing techniques in training stability, pattern preservation accuracy, and generated image fidelity, providing a new theoretical foundation and technical approach for achieving highly stable complex image transformations. Based on this, this invention proposes an unsupervised image transformation method based on Schrödinger bridge theory, such as... Figure 1 As shown, it includes the following steps:

[0044] Step 1: Construct an image domain-to-domain transformation framework based on Schrödinger bridge theory, use a temporal condition generator to simulate the smooth evolution path from the source domain image to the target domain image, and combine it with a discriminator to form adversarial training;

[0045] Step 2: Integrate salient content guidance constraints, global feature consistency constraints, and contrastive learning constraints into the transformation framework to enhance semantic content preservation and detail generation capabilities;

[0046] Step 3: Establish a composite loss function, which combines the Schrödinger bridge path loss, adversarial loss, and various auxiliary constraint losses in a weighted manner to guide model optimization;

[0047] Step 4: Train the temporal condition generator and discriminator by optimizing the algorithm, and after training is completed, use the generator to perform stable and high-fidelity target domain transformation on the source domain image.

[0048] The core of this invention lies in constructing a transformation framework based on Schrödinger's bridge theory and enhancing content preservation and detail generation by integrating multiple auxiliary constraint mechanisms. This method ultimately achieves stable and high-quality image transformation by optimizing a carefully designed composite loss function.

[0049] like Figures 2-4 As shown, the specific implementation process of this invention can be summarized into the following main stages: data preparation and preprocessing, image segmentation, model construction, model training, model weight acquisition, image processing to be converted, image conversion and generation, and image reconstruction and post-processing.

[0050] Figure 2 The overall architecture diagram of the image conversion model shows the connection relationships between the temporal condition generator, discriminator, and various auxiliary constraint modules.

[0051] Figure 3 The network structure diagram of the time condition generator is as follows: It adopts a U-shaped network architecture, displays the input source domain image block and time step information, and outputs the target domain image block after passing through the encoder, decoder and skip-connection.

[0052] Figure 4 The flowchart of the image conversion method in this embodiment includes stages such as data preparation and preprocessing, image segmentation, model training, image conversion, and post-reconstruction processing.

[0053] First, in the data preparation and preprocessing stage, for a specific image domain-to-domain transformation task, source domain image datasets and target domain image datasets need to be collected separately. Crucially, these two datasets do not need to contain paired image samples, fully supporting the unsupervised learning setup. For example, in a medical image virtual staining task, an autofluorescence (AF) microscopic image set can be collected as the source domain, and a hematoxylin-eosin (H&E) staining microscopic image set can be collected as the target domain. After acquisition, necessary preprocessing operations are performed on the raw image data, including but not limited to image format standardization, size adjustment or cropping to fit the model input, and data cleaning to remove low-quality samples.

[0054] The next stage is image segmentation. Given that deep learning models typically have limitations on input image size, and directly processing high-resolution, large images incurs a significant computational burden, this implementation employs an image segmentation strategy. In this stage, the preprocessed source and target domain images are uniformly segmented into fixed-size 256x256 pixel image blocks. These image blocks will serve as the basic input units for model training. It is worth noting that when processing new images to be transformed, the exact same segmentation criteria as in the training stage must be followed.

[0055] Model construction is the core of this invention. Referring to the description in the "Summary of the Invention" section and related framework diagrams, the model mainly consists of a temporal condition generator, a discriminator, and a series of auxiliary constraint modules. The temporal condition generator, based on a mature network architecture, is responsible for simulating the gradual and smooth evolution of an image from the source domain to the target domain. It receives the image patch at the current evolution time step and its corresponding temporal information as input, learning and predicting the dynamics driving the image patch state to evolve towards the target domain. The discriminator adopts an architecture such as PatchGAN, and its responsibility is to distinguish between real target domain image patches and those forged by the generator, thus providing adversarial training signals for the generator. Crucially, the integration of auxiliary constraint modules is essential, including a contrastive learning constraint module to enhance the realism and diversity of local textures, a salient content guidance constraint module to strengthen the preservation and accurate transformation of important content details, and a global feature consistency constraint module to ensure the consistency of multi-scale global semantics and high-level structural features.

[0056] The goal of the model training phase is to optimize the model's parameters. First, a composite loss function is constructed, which integrates the core Schrödinger bridge path loss, adversarial loss, and loss terms calculated by the aforementioned auxiliary constraint modules. During training, batches of data are sampled from the segmented source domain image patch set and the target domain image patch set. For each source domain image patch, a multi-step evolution is performed using a temporal conditional generator to generate pseudo-target domain image patches. Then, these pseudo-target domain image patches and real domain image patches are fed into the discriminator to calculate the adversarial loss. Simultaneously, based on the source image patches and the generated (potentially intermediate or final state) image patches, various auxiliary constraint losses are calculated. Combined with the Schrödinger bridge path loss, a total composite loss value is formed. The Adam optimization algorithm is employed, calculating gradients through backpropagation and alternately updating the network parameters of the temporal conditional generator and the discriminator until the model reaches the convergence criterion or completes the preset training iteration cycle.

[0057] After training is complete, the model weight acquisition phase begins. Once the model training reaches the expected convergence state, the loss function value stabilizes, or the relevant performance metrics no longer show significant improvement on independent validation sets, the trained time-conditional generator network weight parameters are saved. These carefully learned weight parameters solidify the complex mapping relationships and transformation knowledge from the source domain to the target domain that the model has mastered.

[0058] When a new image needs to be transformed into a new domain, the image to be transformed is first processed. For a brand new source image to be transformed, it must first undergo the same preprocessing operations as during the training phase. Then, this preprocessed image is meticulously divided into several independent image patches according to the same size and segmentation method used during training.

[0059] Next comes the core image transformation and generation stage. Each source domain image patch of the image to be transformed, along with its initial temporal information, is input in batches into a temporal conditional generator that has been loaded with previously trained optimized weights. The generator will strictly follow the smooth evolution path learned during the training phase, guided by Schrödinger bridge theory, and through a series of discrete-time step iterations, gradually and meticulously transform each input source domain image patch into its corresponding target domain image patch.

[0060] The final stage is image reconstruction and post-processing. All target domain image blocks independently generated by the temporal condition generator are systematically fused and stitched together based on their precise relative positions in the original image to be converted, thereby reconstructing a complete, large-size target domain image. An overlapping region strategy is employed during the segmentation stage. During stitching, the pixel values ​​of these overlapping regions are processed using a weighted average fusion algorithm to effectively reduce or eliminate potential block boundary effects, ensuring a smoother and more natural visual appearance after stitching. Depending on the specific needs, optional post-processing operations can be performed on the stitched target domain image, such as applying a slight smoothing filter to further eliminate any residual block effects, or performing image enhancement techniques such as contrast enhancement and color correction to comprehensively improve the overall visual quality and subjective impression of the final output image. Through this series of meticulous steps, a stable and high-quality target domain image converted using the method of this invention can be obtained.

[0061] Through the detailed implementation methods described above, the present invention can effectively learn and perform stable and high-quality unsupervised transformations from the source image domain to the target image domain, demonstrating its significant advantages and application potential in handling challenging tasks with large domain differences and high requirements for maintaining complex details.

Claims

1. An unsupervised image conversion imaging method based on the Schrodinger bridge theory, characterized in that: The method comprises the following steps: Step 1: constructing an image domain-to-domain conversion framework based on the Schrödinger bridge theory, using a time-condition generator to simulate a smooth evolution path of source domain images to target domain images, and combining a discriminator to form an adversarial training; Step 2: integrating a salient content guidance constraint, a global feature consistency constraint and a contrastive learning constraint in the conversion framework to enhance the semantic content preservation and detail generation capability; Step 3: establishing a composite loss function, which is a weighted combination of a Schrödinger bridge path loss, an adversarial loss and loss of each auxiliary constraint, to guide model optimization; Step 4: training the time-condition generator and the discriminator through an optimization algorithm, and using the generator to perform stable and high-fidelity target domain conversion on source domain images after training is completed.

2. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 1, characterized in that: The time-condition generator adopts a U-shaped network architecture, takes a source domain image block and an evolution time step as input, and gradually predicts and generates a target domain image block.

3. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 2, characterized in that: The discriminator adopts a PatchGAN structure, discriminates between a target domain real image block and a pseudo target domain image block generated by the generator, and outputs an adversarial training signal.

4. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 1, characterized in that: The saliency content guidance constraint is to compute a consistency loss between the source image and the generated image saliency map , guiding the model to preserve semantically important regions or structures during the transformation.

5. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 1, characterized in that: The global feature consistency constraint uses a deep convolutional neural network VGG16 to extract multi-layer feature maps of the source image and the generated image, and calculates the mean square error as a global feature consistency loss to maintain global semantic and high-level structure feature consistency.

6. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 1, characterized in that: In step S3, the composite loss function is represented as: wherein, for the adversarial loss, the discriminator and the generator are computed, are the parameters of the generator and the discriminator, respectively, and is the intermediate image under the temporal condition, is the distribution of this state, is random noise sampled from a standard normal distribution, is the noise distribution; For the Schr"odinger bridge path loss, used to drive the model learning the optimal stochastic path from the source domain distribution to the target domain distribution, contains the fidelity term and the entropy regularization term; for the target domain image, for the entropy regularization term function, for the backward conditional probability of model learning, , is the weight for the time step; is a contrastive learning loss, is a query feature extracted from a source image at a position ; is a positive sample key feature extracted from a corresponding generated target domain image at the same position ; is a negative sample key feature extracted from other positions or images; is a cosine similarity function, is a temperature coefficient, is a set of layers used for contrast, is a set of spatial positions of the i-th layer. For significant consistency loss, let and be the source domain and generated target domain image samples, respectively, then: where G is a generator from the source domain to the target domain, , is a foreground extraction operator based on a grayscale threshold, is a corresponding threshold, in the form of: For global feature consistency loss,; For the source domain image, To generate an image from the source domain image to the target domain image, This indicates that the VGG16 network is in the... Feature maps output by the layer It is the first The dimension of the layer feature map Let the set of network layers be represented, then: 。 7. The unsupervised image conversion imaging method based on the Schrodinger bridge theory according to claim 1, characterized in that: The optimization algorithm adopts an Adam optimizer.

Citation Information

Cited By

  • Synthetic CT generation method and system, storage medium and electronic equipment

    CN121639863A