Diffusion model based on spherical manifold guidance and training method

By adopting a diffusion model based on spherical manifold guidance in panoramic image generation, combining spherical manifold convolution and spherical manifold guidance module, the spherical distortion problem in panoramic image generation is solved, and high-quality panoramic image generation is achieved.

CN120047567APending Publication Date: 2025-05-27BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510210551.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In panoramic image generation, the prior art has a problem of spherical distortion, especially in the process of converting panoramic content into a flat rectangular format.

Method used

Using a diffusion model based on spherical manifold guidance, the combination of panoramic VQVAE, SMUNet module, panoramic decoder and CLIP module is combined with spherical manifold convolution (SMConv) and spherical manifold guidance module (SME and SMD) to perform forward diffusion and reverse generation in the latent space to solve the spherical distortion problem.

Benefits of technology

Effectively capture the spherical geometric characteristics in panoramic images, reduce the impact of spherical distortion, maintain spatial consistency within the spherical domain, solve the problem of spherical distortion, and improve the quality of the generated panoramic images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047567A_ABST
    Figure CN120047567A_ABST
Patent Text Reader

Abstract

The invention discloses a diffusion model based on spherical manifold guidance and a training method, and the diffusion model based on spherical manifold guidance comprises a panoramic VQVAE, and the panoramic VQVAE comprises an image encoder and a panoramic decoder; the image encoder is used for encoding an input image into a potential space and performing forward diffusion in the potential space to obtain noise data; the SMUNet module is used for de-noising the noise data in reverse generation to obtain recovered data; the panoramic decoder is used for restoring the recovery data into a panoramic image; the CLIP module is used for carrying out text embedding on the SMUNet module in reverse generation; spherical manifold convolution is added into a diffusion model, SMConv is fused into a spherical manifold guiding module, and spherical geometric characteristics in a panoramic image can be optimally captured by the convolution through the SMConv operated on the spherical manifold, so that spherical distortion is solved, and spatial consistency in a spherical domain is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of panoramic image processing, and particularly relates to a diffusion model and a training method based on spherical manifold guidance. Background Art

[0002] Panoramic image generation technology refers to generating panoramic images with a 360° horizontal field of view and a 180° vertical field of view according to text descriptions. This technology has opened up new ways for quickly creating immersive environments and has been widely applied in fields such as virtual reality, augmented reality, and autonomous driving. Although recently introduced deep probability models such as Generative Adversarial Networks (GANs) and diffusion models have played an important role in generating realistic images; the inherent spherical geometric characteristics and distorted scene structures of panoramic images still make panoramic image generation challenging. However, the panoramic image generation method of converting panoramic content into an equirectangular projection (ERP) in a planar rectangular format has the problem of spherical distortion. Summary of the Invention

[0003] In view of this, the purpose of the present invention is to provide a diffusion model and a training method based on spherical manifold guidance to solve the above technical problems.

[0004] To achieve the above purpose, the present invention provides a diffusion model based on spherical manifold guidance, including:

[0005] A panoramic VQVAE, the panoramic VQVAE includes: an image encoder and a panoramic decoder; the image encoder is used to encode the input image into the latent space and perform forward diffusion in the latent space to obtain noise data;

[0006] An SMUNet module, which is used to denoise the noise data in the reverse generation to obtain restored data;

[0007] A panoramic decoder, which is used to restore the restored data to a panoramic image;

[0008] A CLIP module, which is used to perform text embedding on the SMUNet module in the reverse generation.

[0009] As an embodiment of the present invention, the forward diffusion process is as follows:

[0010]

[0011] Where z t represents the noisy spherical feature at time step t, z t-1 represents the noisy spherical feature at time step t - 1, q(zt │z t-1 ) represents the forward diffusion process of gradually adding noise in the latent space, where \(t\in[0, T 0 \), and \(t - 1\in[0, T 0 ;

[0012] The reverse generation process is as follows:

[0013]

[0014] Among them, \(T\) represents the text embedding obtained through the CLIP module, \(\ò ψ (z t , t)\) represents the noise predicted by SMUNet at time step \(t + 1\), \(\psi\) represents the parameters of SMUNet, \(\eta\) represents Gaussian noise, and \(\beta t represents a hyperparameter configured by the noise scheduler, the data after denoising at time step \(t + 1\), represents the noise data obtained after forward diffusion, \(\alpha t = 1 - \beta t .

[0015] As an embodiment of the present invention, the SMUNet module includes:

[0016] Two SPBs, two cyclic encoders, two SMEs, two SMDs, two SMRBs, a cross-attention layer, and a cyclic decoder. The SPB, two cyclic encoders, two SMEs, SMRB, cross-attention layer, SMRB, two SMDs, cyclic decoder, and SPB are connected in sequence.

[0017] As an embodiment of the present invention, the SME includes: two SMRBs, two cross-attention layers, and SMConv. The SMRB, cross-attention layer, SMRB, cross-attention layer, and SMConv are connected in sequence;

[0018] The SMD includes: two SMRBs and two cross-attention layers. The two SMRBs and two cross-attention layers are arranged at intervals and connected in sequence.

[0019] As an embodiment of the present invention, the SMRB module includes: two SMConvs and a time encoding layer. The time encoding layer is arranged between two adjacent SMConvs;

[0020] SMConv performs the following operations:

[0021] The convolution kernel points on the tangent plane \(T p S 2 ​Perform exponential mapping to obtain the mapped convolution kernel point K(m, n), and the exponential mapping is as follows:

[0022]

[0023] Among them, q represents the Cartesian coordinates of the mapped convolution kernel point, represents the tangent point of the tangent plane, (u m , v n ) represents the convolution kernel point on the tangent plane T p S 2 at the predefined position, and and are both the bases of the coordinate system centered at the tangent point p in the tangent plane T p S 2 , v represents a vector,

[0024] Convert the Cartesian coordinates corresponding to the mapped convolution kernel point K(m, n) to spherical coordinates to obtain the spherical coordinate points of the convolution kernel point on the S 2 manifold

[0025] Based on the spherical coordinate points convolve the features on the entire S 2 manifold as follows:

[0026]

[0027] Among them, and are both the features on the S 2 manifold.

[0028] As an embodiment of the present invention, the execution process of the SMRB of the l-th SME block is as follows:

[0029]

[0030] Among them, g ψ (t) represents the learnable time step embedding function, represents the output of this SMRB, represents the output of this SMRB.

[0031] As an embodiment of the present invention, the SPB is an SMConv combined with a 1×1 residual convolution layer.

[0032] On the other hand, the present invention also provides a training method for a diffusion model guided by a spherical manifold, including the following steps:

[0033] Obtain high-quality panoramic VQVAE training data and diffusion model training data;

[0034] Construct an initial panoramic VQVAE, and train the initial panoramic VQVAE based on the panoramic VQVAE training data until convergence to obtain a trained panoramic VQVAE;

[0035] Construct a diffusion model based on the trained panoramic VQVAE and the SMUNet module, and train the diffusion model based on the diffusion model training data until convergence to obtain a trained diffusion model.

[0036] As an embodiment of the present invention, the loss function L for training the panoramic VQVAE VQVAE is as follows:

[0037] L VQVAE = L rec + λ commit L commit + λ per L per

[0038] wherein, L rec represents the reconstruction loss, L commit represents the encoding regularization loss, L per represents the perceptual loss, and λ commit and λ per both represent hyperparameters used to control the weights of different loss function terms during training.

[0039] As an embodiment of the present invention, the loss function L for training the diffusion model Diff is as follows:

[0040] L Diff = λ MSE L MSE + λ SMSE L SMSE

[0041] wherein, λ MSE and λ SMSE both represent hyperparameters, and the calculation formulas of L MSE and L SMSE are as follows:

[0042]

[0043] wherein, the calculation formula of W ij is as follows:

[0044]

[0045] wherein, i represents the index of the width in, and j represents Index of medium height, where W represents width, and H represents height, and ∈ ψ represents the noise predicted by SMUNet, and ∈ represents the Gaussian noise added during training. denotes taking the mathematical expectation with respect to time, training data, and noise, where i ∈ [0, W] and j ∈ [0, H].

[0046] Technical effects of the present invention: During the generation process of panoramic images, by adding Spherical Manifold Convolution (SMConv) to the diffusion model and integrating SMConv into the spherical manifold guidance module (SME and SMD), the SMConv operating on the spherical manifold can optimally capture the spherical geometric features in panoramic images to solve spherical distortion and maintain spatial consistency within the spherical domain.

[0047] Other advantages, objectives, and features of the present invention will be elaborated in the subsequent specification, and to some extent, they are obvious to those skilled in the art, or those skilled in the art can obtain teachings from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To make the objectives, technical solutions, and beneficial effects of the present invention clearer, the present invention provides the following drawings for illustration:

[0049] Figure 1 is the overall network structure diagram of a diffusion model based on spherical manifold guidance according to the present invention;

[0050] Figure 2 is the schematic flow diagram of a training method for a diffusion model based on spherical manifold guidance according to the present invention;

[0051] Figure 3 is the schematic diagram of the partitioned panoramic image quality assessment provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] As Figure 1 shown, the present invention provides a diffusion model based on spherical manifold guidance, including:

[0053] A panoramic VQVAE, where the panoramic VQVAE includes: an image encoder and a panoramic decoder; the image encoder is used to encode the input image into the latent space and perform forward diffusion in the latent space to obtain noise data;

[0054] An SMUNet module, which is used to denoise the noise data during reverse generation to obtain restored data;

[0055] A panoramic decoder for restoring the restored data to a panoramic image;

[0056] A CLIP module for performing text embedding on the SMUNet module during reverse generation;

[0057] The SMUNet module includes:

[0058] Two SPBs, two cyclic encoders, two SMEs, two SMDs, two SMRBs, a cross-attention layer, and a cyclic decoder. The SPB, two cyclic encoders, two SMEs, SMRB, cross-attention layer, SMRB, two SMDs, cyclic decoder, and SPB are connected in sequence.

[0059] The SME includes: two SMRBs, two cross-attention layers, and an SMConv. The SMRB, cross-attention layer, SMRB, cross-attention layer, and SMConv are connected in sequence;

[0060] The SMD includes: two SMRBs and two cross-attention layers. The two SMRBs and two cross-attention layers are arranged at intervals and connected in sequence;

[0061] The SMRB module includes: two SMConvs and a temporal encoding layer. The temporal encoding layer is arranged between two adjacent SMConvs;

[0062] The working principle and beneficial effects of the above technical solution: During the generation of panoramic images, by adding Spherical Manifold Convolution (SMConv) to the diffusion model and integrating SMConv into the spherical manifold guiding modules (SME and SMD), the SMConv operating on the spherical manifold can optimally capture the spherical geometric features in the panoramic image to solve spherical distortion and maintain spatial consistency within the spherical domain; among them, the structure of the SMD is very similar to that of the SME, and the only difference is that its upsampling layer uses nearest-neighbor interpolation; at the same time, the last SME block and the first SMD block do not perform upsampling and downsampling operations; the cyclic encoder and decoder blocks follow the structure in the Latent Diffusion Model (LDM), and all planar convolution operations are replaced by cyclic convolution operations, and the cyclic encoder block uses an average pooling layer for downsampling.

[0063] In one embodiment, the forward diffusion process is as follows:

[0064]

[0065] where z t represents the noise-added spherical feature at time step t, z t-1Denotes the noisy spherical feature at time step t-1, q(z t │z t-1 ) represents the forward diffusion process of gradually adding noise in the latent space, t ∈ [0, T 0 , t-1 ∈ [0, T 0 ;

[0066] The reverse generation process is as follows:

[0067]

[0068] Among them, T represents the text embedding obtained through the CLIP module, ò ψ (z t , t) represents the noise predicted by SMUNet at time step t+1, ψ represents the parameters of SMUNet, η represents Gaussian noise, β t represents a hyperparameter configured by the noise scheduler, The data after denoising at time step t+1, represents the noisy data obtained after forward diffusion, α t = 1-β t ;

[0069] Working principle and beneficial effects of the above technical solution: Based on SMConv, we propose the SMGD model to achieve high-quality panoramic image generation; SMUNet predicts the noise from the noisy spherical features at each time step t and promotes the recovery of the spherical features on the manifold by subtracting the predicted noise; we use spherical processor blocks (SPBs) in the shallower layers of SMUNet, as well as a recurrent encoder and a recurrent decoder based on recurrent convolution; subsequently, a spherical manifold encoder (SME) and a spherical manifold decoder (SMD) are used in the deeper layers, and a module composed of spherical manifold residual blocks (SMRBs) and cross-attention layers is added in the middle layer to integrate text embeddings; the reasons for designing a hybrid architecture combining recurrent convolution and spherical manifold-guided convolution are as follows: First, as the resolution increases, SMConv requires higher video memory overhead compared to traditional convolution. Second, recurrent convolution can correctly process spherical features when combined with spherical manifold convolution. In addition, in the deeper layers of the neural network with lower spatial resolution, the sparse pixel distribution effect near the poles of the panoramic image is significantly enhanced, amplifying the disadvantages of planar convolution. Therefore, the hybrid architecture achieves an optimal trade-off among generating high-quality panoramic images, enhancing computational efficiency, and reducing video memory requirements. Among them, the CLIP module is a pre-trained CLIP ViT / L-14 model for obtaining text embeddings.

[0070] In one embodiment, SMConv performs the following operations:

[0071] Exponentiate the convolution kernel points p S 2 on the tangent plane T to obtain the mapped convolution kernel points K(m, n), and the exponentiation is as follows:

[0072]

[0073] where q represents the Cartesian coordinates of the mapped convolution kernel points, represents the tangent point of the tangent plane, (u m , v n ) represents the predefined position of the convolution kernel points p S 2 on the tangent plane T , and are both on the tangent plane T p S 2The basis of the coordinate system centered at the tangent point p, and v represents a vector.

[0074] Convert the Cartesian coordinates corresponding to the mapping convolution kernel point K(m, n) to spherical coordinates to obtain the spherical coordinate points of the convolution kernel point on the S 2 manifold.

[0075] Based on the spherical coordinate points Perform convolution on the features of the entire S 2 manifold as follows:

[0076]

[0077] Where, and are both features on the S 2 manifold;

[0078] The working principle and beneficial effects of the above technical solution: By directly performing convolution operations in the spherical domain through SMConv, it can effectively maintain the spherical geometric characteristics of panoramic images and reduce the impact of spherical distortion; panoramic images can be regarded as pixels distributed on the S 2 manifold. Through the exponential mapping, the convolution kernel uniformly defined on the tangent plane is projected onto the S 2 manifold, so that the SMConv convolution kernel can be evenly distributed on the entire S 2 manifold;

[0079] Specifically, the S 2 manifold is defined as a set of points x ∈ i 3 in 3D real space with a norm of 1, which can be characterized by spherical coordinates θ ∈ [0, 2π] and 3 ; through the conversion between spherical coordinates and Cartesian coordinates, we can obtain the Cartesian coordinates of points on the sphere and then apply the exponential mapping; specifically, the point on the S 2 manifold can be converted to the corresponding Cartesian coordinate point p = (x, y, z); the shape and position of the convolution kernel are provided by (u p , v 2 ) in T m , v n ; calculate the corresponding Cartesian coordinate position of the mapping convolution kernel point K(m, n) on the S 2 manifold according to the exponential mapping relationship; assume that the center of the convolution kernel K is located at the point p on the S 2 manifold; due to the properties of the exponential mapping, the norm of the mapped point must be 1, thus ensuring that the point remains on the S 2On the manifold; since the position in spherical coordinates is required when performing SMConv, the actual convolution kernel position during the SMConv operation is obtained through the inverse transformation from Cartesian coordinates to spherical coordinates.

[0080] In one embodiment, the execution process of the SMRB of the l-th SME block is as follows:

[0081]

[0082] where g ψ (t) represents the learnable time-step embedding function, represents the output of this SMRB, represents the output of this SMRB;

[0083] SPB is SMConv combined with a 1×1 residual convolutional layer;

[0084] The working principle and beneficial effects of the above technical solution: SPB uses SMConv combined with a 1×1 residual convolutional layer to enhance the conversion performance between different feature channels.

[0085] As Figure 2 shown, the present invention also provides a training method for a diffusion model based on a spherical manifold-guided diffusion model, including the following steps:

[0086] Obtain high-quality panoramic VQVAE training data and diffusion model training data;

[0087] Construct an initial panoramic VQVAE, and train the initial panoramic VQVAE according to the panoramic VQVAE training data until convergence to obtain a trained panoramic VQVAE;

[0088] Based on the trained panoramic VQVAE and the SMUNet module, construct a diffusion model, and train the diffusion model according to the diffusion model training data until convergence to obtain a trained diffusion model;

[0089] The loss function L VQVAE for training the panoramic VQVAE is as follows:

[0090] L VQVAE = L rec + λ commit L commit + λ per L per

[0091] where L rec represents the reconstruction loss, L commit represents the encoding regularization loss, L per represents the perceptual loss, λ commit and λper Both represent hyperparameters used to control the weights of different loss function terms during training;

[0092] The loss function L for training the diffusion model Diff is as follows:

[0093] L Diff = λ MSE L MSE + λ SMSE L SMSE

[0094] where λ MSE and λ SMSE both represent hyperparameters, and the calculation formulas for L MSE and L SMSE are as follows:

[0095]

[0096] where the calculation formula for W ij is as follows:

[0097]

[0098] where i represents the index of the width in , j represents the index of the height in , W represents ψ the width of the noise predicted by SMUNet, ∈ represents the Gaussian noise added during training,

[0099] SMSE can improve the preservation of spherical geometry, especially near the poles of panoramic images, where the sparse pixel distribution of panoramic images is different from that at the equator and has the largest difference from ordinary images.

[0100] Such as Figure 3As shown, the present invention also provides a method for evaluating the quality of panoramic images in partitions. Specifically, the FID score is the de facto standard for evaluating the quality of generated images. However, directly calculating the FID on panoramic images in the rectangular ERP format will introduce distortion. We propose to convert the generated panoramic images from the ERP format to the CMP format composed of 6 square images to calculate the FID score. The mapping from the ERP to the CMP format provides an accurate image transformation mode, converting the rectangular image into a square image, which can be used to evaluate the quality of the generated panoramic images. This conversion reveals the continuity of the left and right boundaries and represents the spherical characteristics by obtaining 4 images representing the regions near the equator and 2 images representing the regions near the poles. Therefore, the 6 square images in the CMP format can be divided into 2 groups, allowing the FID scores to be calculated separately to evaluate the quality of the generated images. Specifically, FID equ is calculated using 4 images near the equator, while FID pole is obtained from 2 images near the poles; FID equ mainly evaluates the quality and visual fidelity of the panoramic content, while FID pole evaluates the degree of preservation of the spherical characteristics of the generated images. Analyzing FID equ and FID pole provides a comprehensive evaluation of the quality of the generated panoramic images.

[0101] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made in form and details without departing from the scope defined by the claims of the present invention.

Claims

1. A diffusion model based on spherical manifold guidance, characterized in that: include: Panoramic VQVAE,Panoramic VQVAE includes: an image encoder and a panoramic decoder; the image encoder is used to encode the input image into a latent space, and perform forward diffusion in the latent space to obtain noise data; The SMUNet module is used to denoise the noisy data in the reverse generation to obtain the restored data; A panoramic decoder, used for restoring the recovered data into a panoramic image; CLIP module for text embedding of SMUNet module in reverse generation.

2. A diffusion model based on spherical manifold guidance according to claim 1, characterized in that: The forward diffusion process is as follows: Among them, z t represents the noisy spherical feature at time step t, z t-1 represents the noisy spherical feature at time step t-1, q(z t │z t-1 ) represents the forward diffusion process of gradually adding noise in the latent space, t∈[0,T0],t-1∈[0,T0]; The reverse generation process is as follows: Where T represents the text embedding obtained by the CLIP module, represents the noise predicted by SMUNet at time step t+1, ψ represents the parameters of SMUNet, η represents Gaussian noise, β t represents a hyperparameter of the noise scheduler configuration, The denoised data at time step t+1, represents the noise data obtained after forward diffusion, α t =1-β t .

3. The diffusion model based on spherical manifold guidance according to claim 1, characterized in that: SMUNet modules include: Two SPBs, two recurrent encoders, two SMEs, two SMDs, two SMRBs, a criss-cross attention layer and a recurrent decoder, SPBs, two recurrent encoders, two SMEs, SMRBs, a criss-cross attention layer, SMRBs, two SMDs, a recurrent decoder and SPBs are connected sequentially.

4. A diffusion model based on spherical manifold guidance according to claim 3, characterized in that: SME includes: two SMRBs, two cross-attention layers and SMConv, SMRB, cross-attention layer, SMRB, cross-attention layer and SMConv are connected in sequence; SMD includes: two SMRBs and two cross-attention layers, and the two SMRBs and the two cross-attention layers are arranged at intervals and connected in sequence.

5. The diffusion model based on spherical manifold guidance according to claim 3, characterized in that: The SMRB module includes: two SMConvs and a time coding layer, wherein the time coding layer is arranged between two adjacent SMConvs; SMConv execution includes the following operations: The tangent plane T p S 2 The convolution kernel points on Perform exponential mapping to obtain the mapped convolution kernel point K(m, n). The exponential mapping is as follows: Among them, q represents the Cartesian coordinates of the mapped convolution kernel point, represents the tangent point of the tangent plane, (u m , v n ) represents the tangent plane T q S 2 Convolution kernel point A predefined location on and The tangent plane T p S 2 The basis of the coordinate system centered at the tangent point p, v represents a vector, Convert the Cartesian coordinates corresponding to the mapped convolution kernel point K(m, n) to spherical coordinates to obtain the convolution kernel point in S 2 Spherical coordinate points on a manifold Based on spherical coordinates For the entire S 2 The features on the manifold are convolved as follows: in, and All S 2 Features on the manifold.

6. A diffusion model based on spherical manifold guidance according to claim 4, characterized in that: The execution process of SMRB for the lth SME block is as follows: Among them, g ψ (t) represents the learnable time-step embedding function, Represents the output of the SMRB, Represents the output of this SMRB.

7. The diffusion model based on spherical manifold guidance according to claim 3, characterized in that: SPB is SMConv combined with a 1×1 residual convolution layer.

8. A method for training a diffusion model based on spherical manifold guidance, characterized in that: The following steps are involved: Obtain high-quality panoramic VQVAE training data and diffusion model training data; Construct an initial panoramic VQVAE, train the initial panoramic VQVAE according to the panoramic VQVAE training data until convergence, and obtain a trained panoramic VQVAE; A diffusion model is constructed based on the trained panoramic VQVAE and SMUNet modules, and the diffusion model is trained until convergence according to the diffusion model training data to obtain a trained diffusion model.

9. The method for training a diffusion model based on spherical manifold guidance according to claim 8, characterized in that: The loss function L for training panoramic VQVAE VQVAE , as shown below: L VQVAE =L rec +λ commit L commit +λ per L per Among them, L rec represents the reconstruction loss, L commit represents the encoding regularization loss, L per represents the perceptual loss, λ commit and λ per They all represent hyperparameters, which are used to control the weights of different loss function terms during training.

10. The method for training a diffusion model based on spherical manifold guidance according to claim 8, characterized in that: The loss function L for training the diffusion model Diff , as shown below: L Diff =λ MSE L MSE +λ SMSE L SMSE Among them, λ MSE and λ SMSE Both represent hyperparameters, L MSE and L SMSE The calculation formula is as follows: Among them, W ij The calculation formula is as follows: Among them, i represents The index of the middle width, j represents The index of medium height, W represents The width of H represents The height of ψ represents the noise predicted by SMUNet, ∈ represents the Gaussian noise added during training, It represents the mathematical expectation of time, training data, and noise, i∈[0,W], j∈[0,H].